Articles | Open Access | DOI: https://doi.org/10.55640/ijcsis/Volume11Issue09-03

Secure and Trustworthy LLM-Based Code Generation for Enterprise Software Systems

arjun Mehta , Department of Computer Science and Engineering, National Institute of Software Technology, Bengaluru, India

Abstract

Large Language Models (LLMs) have introduced a new paradigm for software development by enabling developers to generate, transform, explain, and test source code through natural-language interaction. Their adoption in enterprise environments, however, introduces significant concerns regarding reliability, verification, security, maintainability, and confidence in generated artifacts. Enterprise software operates under stringent requirements for correctness, continuous integration, traceability, and operational stability, making uncontrolled reliance on automatically generated code problematic. This research-and-review paper develops a conceptual framework for secure and trustworthy LLM-based code generation by integrating LLM-assisted development with established principles of software testing, automated verification, static and dynamic program analysis, continuous integration, and flaky-test management. The methodology synthesizes the provided literature to identify how testing reliability, failure diagnosis, test prioritization, and large-scale continuous testing can serve as complementary safeguards for LLM-generated code. The analysis indicates that trustworthy code generation should not be treated as a generation-only problem; rather, it should be implemented as a closed-loop process consisting of generation, validation, testing, diagnosis, risk assessment, and controlled integration. Particular attention is given to flaky tests because nondeterministic validation can undermine confidence in otherwise correct code. The resulting framework emphasizes layered verification, automated quality gates, explainable validation evidence, and human oversight for high-risk enterprise components. The study contributes a research-oriented model for integrating LLM capabilities with mature software assurance mechanisms and identifies directions for future empirical evaluation.

Keywords

Large Language Models, Code Generation, Software Security, Trustworthy AI

References

“Facebook testing and verification request for proposals,” Accessed: Feb. 4, 2023. [Online]. Available: https://research.fb.com/programs/research-awards/proposals/facebook-testing-and-verification-request-for-proposals-2019

“Test verification,” Accessed: Feb. 4, 2023. [Online]. Available: https://developer.mozilla.org/en-US/docs/Mozilla/QA/Test_Verification

“TotT: Avoiding flakey tests,” Google. Accessed: Feb. 4, 2023. [Online]. Available: http://googletesting.blogspot.com/2008/04/tott-avoiding-flakey-tests.html

M. Eck, F. Palomba, M. Castelluccio, and A. Bacchelli, “Understanding flaky tests: The developer's perspective,” in Proc. 27th Joint Meeting Found. Softw. Eng., 2019, pp. 830–840.

M. Harman and P. O’Hearn, “From start-ups to scale-ups: Opportunities and open problems for static and dynamic program analysis,” in Proc. 18th Int. Work. Conf. Source Code Anal. Manipulation (SCAM), 2018, pp. 1–23.

K. Herzig and N. Nagappan, “Empirically detecting false test alarms using association rules,” in Proc. 37th Int. Conf. Softw. Eng. (ICSE), 2015, pp. 39–48.

K. Herzig, M. Greiler, J. Czerwonka, and B. Murphy, “The art of testing less without sacrificing quality,” in Proc. 37th Int. Conf. Softw. Eng. (ICSE), 2015, pp. 483–493.

H. Jiang, X. Li, Z. Yang, and J. Xuan, “What causes my test alarm? Automatic cause analysis for test alarms in system and integration testing,” in Proc. 39th Int. Conf. Softw. Eng. (ICSE), 2017, pp. 712–723.

E. Kowalczyk, K. Nair, Z. Gao, L. Silberstein, T. Long, and A. Memon, “Modeling and ranking flaky tests at Apple,” in Proc. 42nd Int. Conf. Softw. Eng., Softw. Eng. Pract. (ICSE SEIP), 2020, pp. 110–119.

W. Lam, P. Godefroid, S. Nath, A. Santhiar, and S. Thummalapenta, “Root causing flaky tests in a large-scale industrial setting,” in Proc. Int. Symp. Softw. Testing Anal. (ISSTA), 2019, pp. 101–111.

W. Lam, K. Muşlu, H. Sajnani, and S. Thummalapenta, “A study on the lifecycle of flaky tests,” in Proc. 42nd Int. Conf. Softw. Eng. (ICSE), 2020, pp. 1471–1482.

T. Leesatapornwongsa, X. Ren, and S. Nath, “FlakeRepro: Automated and efficient reproduction of concurrency-related flaky tests,” in Proc. 30th Joint Meeting Found. Softw. Eng. (ESEC/FSE), 2022, pp. 1509–1520.

Q. Luo, F. Hariri, L. Eloussi, and D. Marinov, “An empirical analysis of flaky tests,” in Proc. ACM SIGSOFT 22nd Symp. Found. Softw. Eng., 2014, pp. 643–653.

J. Malm, A. Causevic, B. Lisper, and S. Eldh, “Automated analysis of flakiness-mitigating delays,” in Proc. 1st Int. Conf. Automat. Softw. Test (AST), 2020, pp. 81–84.

A. Memon, “Taming Google-scale continuous testing,” in Proc. 39th Int. Conf. Softw. Eng., Softw. Eng. Pract. Track (ICSE SEIP), 2017, pp. 233–242.

J. Micco, “The state of continuous integration testing at Google,” in Proc. 10th Int. Conf. Softw. Testing, Verification Validation Keynote (ICST), 2017.

O. Parry, G. M. Kapfhammer, M. Hilton, and P. McMinn, “A survey of flaky tests,” Trans. Softw. Eng. Methodol., vol. 31, no. 1, pp. 1–74, 2021.

M. T. Rahman and P. C. Rigby, “The impact of failing, flaky, and high failure tests on the number of crash reports associated with Firefox builds,” in Proc. 26th Joint Meeting Found. Softw. Eng. (ESEC/FSE), 2018, pp. 857–862.

M. H. U. Rehman and P. C. Rigby, “Quantifying no-fault-found test failures to prioritize inspection of flaky tests at Ericsson,” in Proc. 29th Joint Meeting Found. Softw. Eng., Ind. Track (ESEC/FSE), 2021.

C. Ziftci and J. Reardon, “Who broke the build?: Automatically identifying changes that induce test failures in continuous integration at Google scale,” in Proc. 39th Int. Conf. Softw. Eng. (ICSE), 2017, pp. 113–122.

Kongari, S. S. R., Kumar, S. K.., & Kumar, A. (2026). Trustworthy and Secure LLM-Assisted Code Generation for Enterprise Software Development. International Journal of Data Science and Machine Learning, 6(01), 222-237. https://doi.org/10.55640/ijdsml-06-01-04

Article Statistics

Downloads

Download data is not yet available.

Copyright License

Download Citations

How to Cite

arjun Mehta. (2026). Secure and Trustworthy LLM-Based Code Generation for Enterprise Software Systems. International Journal of Computer Science & Information System, 11(09), 19–26. https://doi.org/10.55640/ijcsis/Volume11Issue09-03