Abstract
The rapid advancement of large language models (LLMs) has significantly influenced various domains, including software engineering, where their potential to transform raditional practices is increasingly recognized. This systematic literature review investigates the role of LLMs in software requirements engineering, aiming to synthesize existing research, identify key trends, and highlight gaps in the current knowledge. We examine the multifaceted applications of LLMs across several dimensions, such as requirements elicitation, analysis, and specification, as well as their broader implications in code generation, testing, and security. The methodology involves a rigorous selection and analysis of relevant studies, followed by a thematic categorization to provide a comprehensive overview of the field. Our findings reveal that LLMs demonstrate considerable promise in automating and improving the accuracy of requirements-related tasks, yet challenges such as model interpretability, ethical concerns, and integration with existing workflows remain unresolved. The review also underscores the growing emphasis on prompt engineering and fine-tuning techniques to adapt LLMs for domain-specific needs. While the adoption of LLMs in software requirements engineering is still evolving, the evidence suggests a paradigm shift toward more intelligent and efficient practices. We conclude by discussing future research directions, emphasizing the need for empirical validation, interdisciplinary collaboration, and the development of robust frameworks to address the limitations and risks associated with LLM deployment in this critical area of software engineering.
Keywords
References
Akbar M. A., Khan A. A., & Liang P. (2023). Ethical aspects of ChatGPT in software engineering research. IEEE Transactions on Artificial Intelligence, Vol. 6, Issue 2, pp. 254-267. Article
Alagarsamy S., Tantithamthavorn C., Takerngsaksiri W., Arora C., & Aleti A. (2025). Enhancing large language models for text-to-testcase generation. Journal of Systems and Software, Vol. 230, 112531. Article
Alshahwan N., Chheda J., Finogenova A., Gokkaya B., Harman M., Harper, I., ... & Wang, E. (2024, July). Automated unit test improvement using large language models at meta. 32nd ACM International Conference on the Foundations of Software Engineering, pp. 185-196. Article
Arora C., Grundy J., & Abdelrazek M. (2024). Advancing requirements engineering through generative ai: Assessing the role of LLMs. Generative AI for Effective Software Development, pp. 129-148. Cham: Springer Nature Switzerland. Article
Aytekin M. C., Yılmaz F. G., & Demirezen M. U. (2026). Automating code generation for a new ecosystem: establishing baselines with large language model-based code generation for ArkTS and HarmonyOS. Automated Software Engineering, Vol. 33, Issue 2. Article
Beg A., O’Donoghue D., & Monahan R. (2025). Formalising Software Requirements with Large Language Models. ADAPT Annual Conference.
Bjarnason E., Unterkalmsteiner M., Borg M., & Engström E. (2016). A multi-case study of agile requirements engineering and the use of test cases as requirements. Information and Software Technology, Vol. 77, pp. 61-79. Article
Bouzenia I., & Pradel M. (2025). You name it, i run it: An LLM agent to execute tests of arbitrary projects. ACM on Software Engineering, Vol. 2, pp. 1054-1076. Article
Bui T. L., Dam H. K., & Hoda R. (2025, November). An LLM-based multi-agent framework for agile effort estimation. 40th IEEE/ACM International Conference on Automated Software Engineering, pp. 1032-1043. Article
Chen X., Gao C., Chen C., Zhang G., & Liu Y. (2025). An empirical study on challenges for LLM application developers. ACM Transactions on Software Engineering and Methodology, Vol. 34, Issue 7, pp. 1-37. Article
Dalpiaz F., & Niu N. (2020). Requirements engineering in the days of artificial intelligence. IEEE Software, Vol. 37, Issue 4, pp. 7-10. Article
Daun M., & Brings J. (2023, June). How ChatGPT will change software engineering education. In Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education, Volume 1, pp. 110-116. Article
Deng Y., Xia C. S., Yang C., Zhang S. D., Yang S., & Zhang L. (2024). Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries. 46th IEEE/ACM International Conference on Software Engineering, pp. 1-13. Article
Dong Y., Jiang X., Jin Z., & Li G. (2024). Self-collaboration code generation via ChatGpt. ACM Transactions on Software Engineering and Methodology, Vol. 33, Issue 7, pp. 1-38. Article
Doris A. C., Grandi D., Tomich R., Alam M. F., Ataei M., Cheong H., & Ahmed F. (2025). Designqa: A multimodal benchmark for evaluating large language models’ understanding of engineering documentation. Journal of Computing and Information Science in Engineering, Vol. 25, Issue 2, 021009. Article
Elghariani K., & Kama N. (2016). Review on Agile requirements engineering challenges. 3rd International Conference on Computer and Information Sciences, pp. 507-512. Article
Ezzini S., Abualhaija S., Arora C., & Sabetzadeh M. (2023). Ai-based question answering assistance for analyzing natural-language requirements. IEEE / ACM 45th International Conference on Software Engineering, pp. 1277-1289. Article
Fan Z., Gao X., Mirchev M., Roychoudhury A., & Tan S. H. (2023). Automated repair of programs from large language models. IEEE/ACM 45th International Conference on Software Engineering, pp. 1469-1481. Article
Fatima S., Ghaleb T. A., & Briand L. (2022). Flakify: A black-box, language model-based predictor for flaky tests. IEEE Transactions on Software Engineering, Vol. 49, Issue 4, pp. 1912-1927. Article
Feldt R., Kang S., Yoon J., & Yoo S. (2023). Towards autonomous testing agents via conversational large language models. 38th IEEE/ACM International Conference on Automated Software Engineering, pp. 1688-1693. IEEE. Article
Fu M., Tantithamthavorn C. K., Nguyen V., & Le T. (2023). ChatGpt for vulnerability detection, classification, and repair: How far are we? 30th Asia-Pacific Software Engineering Conference, pp. 632-636. Article
Guo Y., Patsakis C., Hu Q., Tang Q., & Casino F. (2024). Outside the comfort zone: Analysing llm capabilities in software vulnerability detection. European Symposium on Research in Computer Security, pp. 271-289. Cham: Springer Nature Switzerland. Article
Han H., Kim J., Yoo J., Lee Y., & Hwang S. W. (2024). Archcode: Incorporating software requirements in code generation with large language models. Annual Meeting of the Association for Computational Linguistics, Vol. 1, pp. 13520-13552. Article
Heyn H. M., Knauss E., Muhammad A. P., Eriksson O., Linder J., Subbiah P., ... & Tungal S. (2021, May). Requirement engineering challenges for ai-intense systems development. IEEE/ACM 1st Workshop on AI Engineering-Software Engineering for AI, pp. 89-96. IEEE. Article
Jain N., Vaidyanath S., Iyer A., Natarajan N., Parthasarathy S., Rajamani S., & Sharma R. (2022). Jigsaw: Large language models meet program synthesis. 44th International Conference on Software Engineering, pp. 1219-1231. Article
Jain N., Vaidyanath S., Iyer A., Natarajan N., Parthasarathy S., Rajamani S., & Sharma R. (2022). Jigsaw: Large language models meet program synthesis. 44th International Conference on Software Engineering, pp. 1219-1231. Article
Jimenez C. E., Yang J., Wettig A., Yao S., Pei K., Press O., & Narasimhan K. (2024, May). Swe-bench: Can language models resolve real-world github issues? International Conference on Learning Representations, Vol. 2024, pp. 54107-54157.
Jin M., Shahriar S., Tufano M., Shi X., Lu S., Sundaresan N., & Svyatkovskiy A. (2023). Inferfix: End-to-end program repair with LLMs. 31st ACM Joint European Software Engineering Conference and Symposium on The Foundations of Software Engineering, pp. 1646-1656. Article
Kaddour J., Harris J., Mozes M., Bradley H., Raileanu R., & McHardy R. (2023). Challenges and applications of large language models. arXiv preprint arXiv:2307.10169. https://doi.org/10.48550/arXiv.2307.10169 Article
Kang S., Yoon J., & Yoo S. (2023, May). Large language models are few-shot testers: Exploring LLM-based general bug reproduction. IEEE/ACM 45th International Conference on Software Engineering, pp. 2312-2323. Article
Khare A., Dutta S., Li Z., Solko-Breslin A., Alur R., & Naik M. (2025). Understanding the effectiveness of large language models in detecting security vulnerabilities. IEEE Conference on Software Testing, Verification and Validation, pp. 103-114. Article
Khattab O., Singhvi A., Maheshwari P., Zhang Z., Santhanam K., Vardhamanan S., ... & Potts C. (2023). Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714. Article
Khojah R., Mohamad M., Leitner P., & de Oliveira Neto F. G. (2024). Beyond code generation: An observational study of chatgpt usage in software engineering practice. ACM on Software Engineering, Vol. 1(FSE), pp. 1819-1840. Article
Kozov V., Ivanova G., & Atanasova D. (2024). Practical Application of AI and Large Language Models in Software Engineering Education. International Journal of Advanced Computer Science & Applications, Vol. 15, Issue 1. Article
Krishna M., Gaur B., Verma A., & Jalote P. (2024, June). Using LLMs in software requirements specifications: An empirical evaluation. 32nd International Requirements Engineering Conference, pp. 475-483. Article
Li J., Li G., Li Y., & Jin Z. (2025). Structured chain-of-thought prompting for code generation. ACM Transactions on Software Engineering and Methodology, Vol. 34, Issue 2, pp. 1-23. Article
Li J., Rabbi F., Cheng C., Sangalay A., Tian Y., & Yang J. (2026). An exploratory study on fine-tuning large language models for secure code generation. Empirical Software Engineering, Vol. 31, Issue 4. Article
Li J., Zhao Y., Li Y., Li G., & Jin Z. (2024). Acecoder: An effective prompting technique specialized in code generation. ACM Transactions on Software Engineering and Methodology, Vol. 33, Issue 8, pp. 1-26. Article
Liu J., Xia C. S., Wang Y., & Zhang L. (2023). Is your code generated by ChatGpt really, correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems, Vol. 36, pp. 21558-21572. Article
Lops A., Narducci F., Ragone A., Trizio M., & Bartolini C. (2025, March). A system for automated unit test generation using large language models and assessment of generated test suites. IEEE International Conference on Software Testing, Verification and Validation Workshops, pp. 29-36. Article
Lu Q., Zhu L., Xu X., Xing Z., Harrer S., & Whittle J. (2024). Towards responsible generative ai: A reference architecture for designing foundation model-based agents. IEEE 21st International Conference on Software Architecture Companion, pp. 119-126. Article
Luitel D., Hassani S., & Sabetzadeh M. (2023). Using language models for enhancing the completeness of natural-language requirements. In International Working Conference on Requirements Engineering: Foundation for Software Quality, pp. 87-104. Cham: Springer Nature Switzerland. Article
Manish S. (2024). An autonomous multi-agent LLM framework for agile software development. International Journal of Trend in Scientific Research and Development, Vol. 8, Issue 5, pp. 892-898.Article
Mathews N. S., & Nagappan M. (2024, October). Test-driven development and llm-based code generation. 39th IEEE/ACM International Conference on Automated Software Engineering, pp. 1583-1594. Article
Minaee S., Mikolov T., Nikzad N., Chenaghlu M., Socher R., Amatriain X., & Gao J. (2024). Large language models: A survey. arXiv preprint arXiv:2402.06196. Article
Molina F., Gorla A., & d’Amorim M. (2025). Test Oracle Automation in the era of LLMs. ACM Transactions on Software Engineering and Methodology, Vol. 34, Issue 5, pp. 1-24. Article
Mu F., Shi L., Wang S., Yu Z., Zhang B., Wang C., ... & Wang Q. (2024). Clarifygpt: A framework for enhancing llm-based code generation via requirements clarification. ACM on Software Engineering, 1, pp. 2332-2354. Article
Nam D., Macvean A., Hellendoorn V., Vasilescu B., & Myers B. (2024, April). Using an LLM to help with code understanding. IEEE/ACM 46th International Conference on Software Engineering, pp. 1-13. Article
Neelapu M. Enhancing Software Testing Efficiency with Generative AI and Large Language Models. IJLRP-International Journal of Leading Research Publication, Vol. 5, Issue 12. Article
Noever D. (2023). Can large language models find and fix vulnerable software? arXiv preprint arXiv:2308.10345. Article
Nong Y., Aldeen M., Cheng L., Hu H., Chen F., & Cai H. (2024). Chain-of-thought prompting of large language models for discovering and fixing software vulnerabilities. arXiv preprint arXiv:2402.17230. Article
Ozkaya I. (2023). Application of large language models to software engineering tasks: Opportunities, risks, and implications. IEEE software, Vol. 40, Issue 3, pp. 4-8. Article
Page M. J., McKenzie J. E., Bossuyt P. M., Boutron I., Hoffmann T. C., Mulrow C. D., ... & Moher D. (2021). The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. bmj, Vol. 372. Article
Plein L., Ouédraogo W. C., Klein J., & Bissyandé T. F. (2024). Automatic generation of test cases based on bug reports: a feasibility study with large language models. IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, pp. 360-361. Article
Pudari R., & Ernst N. A. (2023). From copilot to pilot: Towards AI supported software development. arXiv preprint arXiv:2303.04142. Article
Rajbhoj A., Somase A., Kulkarni P., & Kulkarni V. (2024). Accelerating software development using generative AI: ChatGPT case study. 17th Innovations in Software Engineering Conference, pp. 1-11. Article
Rao N., Jain K., Alon U., Le Goues C., & Hellendoorn V. J. (2023). CAT-LM training language models on aligned code and tests. 38th IEEE/ACM International Conference on Automated Software Engineering, pp. 409-420. Article
Rodriguez A. D., Dearstyne K. R., & Cleland-Huang J. (2023). Prompts matter: Insights and strategies for prompt engineering in automated software traceability. 31st International Requirements Engineering Conference Workshops, pp. 455-464. Article
Rodriguez-Cardenas D. (2025, June). Towards More Interpretable Large Language Models for Code. 33rd ACM International Conference on the Foundations of Software Engineering, pp. 1270-1272. Article
Ronanki K., Berger C., & Horkoff J. (2023). Investigating chatgpt’s potential to assist in requirements elicitation processes. 49th IEEE Euromicro Conference on Software Engineering and Advanced Applications, pp. 354-361. Article
Ronanki K., Cabrero-Daniel B., Horkoff J., & Berger C. (2024). Requirements engineering using generative ai: Prompts and prompting patterns. Generative AI for Effective Software Development, pp. 109-127. Cham: Springer Nature Switzerland. Article
Ross S. I., Martinez F., Houde S., Muller M., & Weisz J. D. (2023, March). The programmer’s assistant: Conversational interaction with a large language model for software development. 28th International Conference on Intelligent User Interfaces, pp. 491-514. Article
Sallou J., Durieux T., & Panichella A. (2024, April). Breaking the silence: the threats of usingLLMs in software engineering. ACM/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results, pp. 102-106. Article
Santos R., Santos I., Magalhaes C., & de Souza Santos R. (2024). Are we testing or being tested? exploring the practical applications of large language models in software testing. IEEE Conference on Software Testing, Verification and Validation, pp. 353-360. Article
Shin J., Tang C., Mohati T., Nayebi M., Wang S., & Hemmati H. (2023). Prompt engineering or fine tuning: An empirical assessment of large language models in automated software engineering tasks. arXiv preprint arXiv:2310.10508. Article
Shirafuji A., Oda Y., Suzuki J., Morishita M., & Watanobe Y. (2023, December). Refactoring programs using large language models with few-shot examples. 30th Asia-Pacific Software Engineering Conference, pp. 151-160. Article
Siddiq M. L., Da Silva Santos J. C., Tanvir R. H., Ulfat N., Al Rifat F., & Carvalho Lopes V. (2024, June). Using large language models to generate junit tests: An empirical study. 28th International Conference on Evaluation and Assessment in Software Engineering, pp. 313-322. Article
Singla T., Anandayuvaraj D., Kalu K. G., Schorlemmer T. R., & Davis J. C. (2023). An empirical study on using large language models to analyze software supply chain security failures. Workshop on Software Supply Chain Offensive Research and Ecosystem Defenses, pp. 5-15. Article
Son H. S. (2025). Using large language models in software requirements analysis. Edelweiss Appl. Sci. Technol, Vol. 9, Issue 3, pp. 856-863.
Subahi, A. F. (2026). Prompt Engineering for Advancing Software Engineering: Using Generative Language Models in the Development of Domain-Specific Languages. IEEE Access, Vol. 14, pp. 22834-22850. Article
Sun W., Miao Y., Li Y., Zhang H., Fang C., Liu Y., ... & Chen Z. (2025). Source code summarization in the era of large language models. IEEE/ACM 47th International Conference on Software Engineering, pp. 1882-1894. Article
Tihanyi N., Bisztray T., Jain R., Ferrag M. A., Cordeiro L. C., & Mavroeidis V. (2023). The formai dataset: Generative ai in software security through the lens of formal verification. 19th International Conference on Predictive Models and Data Analytics in Software Engineering, pp. 33-43. Article
Tihanyi N., Charalambous Y., Jain R., Ferrag M. A., & Cordeiro L. C. (2025). A new era in software security: Towards self-healing software via large language models and formal verification. IEEE/ACM International Conference on Automation of Software Test, pp. 136-147. IEEE. Article
Treude C., & Hata H. (2023). She elicits requirements and he tests: Software engineering gender bias in large language models. IEEE/ACM 20th International Conference on Mining Software Repositories, pp. 624-629. Article
Wang Y., Kordi Y., Mishra S., Liu A., Smith N. A., Khashabi D., & Hajishirzi H. (2023). Self-instruct: Aligning language models with self-generated instructions. 61st Annual Meeting of the Association for Computational Linguistics, Vol. 1, pp. 13484-13508. Article
Wang Z., Liu K., Li G., & Jin Z. (2024). Hits: High-coverage llm-based unit test generation via method slicing. 39th IEEE/ACM International Conference on Automated Software Engineering, pp. 1258-1268. Article
Wei B. (2024). Requirements are all you need: From requirements to code with LLM. In 2024 IEEE 32nd International Requirements Engineering Conference, pp. 416-422. Article
White J., Hays S., Fu Q., Spencer-Smith J., & Schmidt D. C. (2024). Chatgpt prompt patterns for improving code quality, refactoring, requirements elicitation, and software design. In Generative ai for effective software development, pp. 71-108. Cham: Springer Nature Switzerland. Article
Xia C. S., Paltenghi M., Le Tian J., Pradel M., & Zhang L. (2024). Fuzz4all: Universal fuzzing with large language models. IEEE/ACM 46th International Conference on Software Engineering, pp. 1-13. Article
Xue Z., Li L., Tian S., Chen X., Li P., Chen L., ... & Zhang M. (2024). Llm4fin: Fully automating LLM-powered test case generation for fintech software acceptance testing. 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, pp. 1643-1655. Article
Yang A. Z., Le Goues C., Martins R., & Hellendoorn V. (2024). Large language models for test-free fault localization. 46th IEEE/ACM International Conference on Software Engineering, pp. 1-12. Article
Yang Z., Liu F., Yu Z., Keung J. W., Li J., Liu S., ... & Li G. (2024). Exploring and unleashing the power of large language models in automated code translation. ACM on Software Engineering, Vol. 1(FSE), pp. 1585-1608. Article
Yu S., Fang C., Ling Y., Wu C., & Chen Z. (2023). LLM for test script generation and migration: Challenges, capabilities, and opportunities. IEEE 23rd International Conference on Software Quality, Reliability, and Security, pp. 206-217. Article
Zhang Y., Ruan H., Fan Z., & Roychoudhury A. (2024, September). Autocoderover: Autonomous program improvement. 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, pp. 1592-1604. Article
Zhang Y., Xie Y., Lit S., Liu K., Wang C., Jia Z., ... & Liao X. (2025, April). Unseen horizons: Unveiling the real capability of LLM code generation beyond the familiar. IEEE/ACM 47th International Conference on Software Engineering, pp. 604-615. Article
Zhang Z., Rayhan M., Herda T., Goisauf M., & Abrahamsson P. (2024). Llm-based agents for automating the enhancement of user story quality: An early report. International Conference on Agile Software Development, pp. 117-126. Cham: Springer Nature Switzerland. Article
Zheng Z., Ning K., Zhong Q., Chen J., Chen W., Guo L., ... & Wang Y. (2025). Towards an understanding of large language models in software engineering tasks. Empirical Software Engineering, Vol. 30, Issue 2. Article
Zhou X., Zhang T., & Lo D. (2024). Large language model for vulnerability detection: Emerging results and future directions. ACM/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results, pp. 47-51. Article
Zhu K., Wang J., Zhou J., Wang Z., Chen H., Wang Y., & Xie X. (2023). Promptrobust: Towards evaluating the robustness of large language models on adversarial prompts. ACM Workshop on Large AI Systems and Models with Privacy and Safety Analysis, pp. 57-68. Article
Ziegler D. M., Stiennon N., Wu J., Brown T. B., Radford A., Amodei D., ... & Irving G. (2019). Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593. Article