Authors
Shafaque Aziz
Department of Computer Science and Engineering, Government Engineering College, Samastipur
Bihar, India
Khushboo
Amity Institute of Information Technology, Amity University, Patna, Bihar, India
Saim Akhtar
Department of Computer Engineering, J.C. Bose University of Science and Technology, YMCA Faridabad, Haryana, India
Sameera Jawed
Department of Computer Science and Technology, Maharaja Agrasen Institute of Technology Affiliated with Guru Gobind Singh Indraprastha University, New Delhi, India

Abstract

The rapid advancement of large language models (LLMs) has significantly influenced various domains, including software engineering, where their potential to transform raditional practices is increasingly recognized. This systematic literature review investigates the role of LLMs in software requirements engineering, aiming to synthesize existing research, identify key trends, and highlight gaps in the current knowledge. We examine the multifaceted applications of LLMs across several dimensions, such as requirements elicitation, analysis, and specification, as well as their broader implications in code generation, testing, and security. The methodology involves a rigorous selection and analysis of relevant studies, followed by a thematic categorization to provide a comprehensive overview of the field. Our findings reveal that LLMs demonstrate considerable promise in automating and improving the accuracy of requirements-related tasks, yet challenges such as model interpretability, ethical concerns, and integration with existing workflows remain unresolved. The review also underscores the growing emphasis on prompt engineering and fine-tuning techniques to adapt LLMs for domain-specific needs. While the adoption of LLMs in software requirements engineering is still evolving, the evidence suggests a paradigm shift toward more intelligent and efficient practices. We conclude by discussing future research directions, emphasizing the need for empirical validation, interdisciplinary collaboration, and the development of robust frameworks to address the limitations and risks associated with LLM deployment in this critical area of software engineering.

Keywords

Software Engineering Requirements Engineering Large Language Model Systematic Literature Review Approximation spaces Requirements Prioritization

References

    Akbar M. A., Khan A. A., & Liang P. (2023). Ethical aspects of ChatGPT in software engineering research. IEEE Transactions on Artificial Intelligence, Vol. 6, Issue 2, pp. 254-267. Article

    Alagarsamy S., Tantithamthavorn C., Takerngsaksiri W., Arora C., & Aleti A. (2025). Enhancing large language models for text-to-testcase generation. Journal of Systems and Software, Vol. 230, 112531. Article

    Alshahwan N., Chheda J., Finogenova A., Gokkaya B., Harman M., Harper, I., ... & Wang, E. (2024, July). Automated unit test improvement using large language models at meta. 32nd ACM International Conference on the Foundations of Software Engineering, pp. 185-196. Article

    Arora C., Grundy J., & Abdelrazek M. (2024). Advancing requirements engineering through generative ai: Assessing the role of LLMs. Generative AI for Effective Software Development, pp. 129-148. Cham: Springer Nature Switzerland. Article

    Aytekin M. C., Yılmaz F. G., & Demirezen M. U. (2026). Automating code generation for a new ecosystem: establishing baselines with large language model-based code generation for ArkTS and HarmonyOS. Automated Software Engineering, Vol. 33, Issue 2. Article

    Beg A., O’Donoghue D., & Monahan R. (2025). Formalising Software Requirements with Large Language Models. ADAPT Annual Conference. Bjarnason E., Unterkalmsteiner M., Borg M., & Engström E. (2016). A multi-case study of agile requirements engineering and the use of test cases as requirements. Information and Software Technology, Vol. 77, pp. 61-79. Article

    Bouzenia I., & Pradel M. (2025). You name it, i run it: An LLM agent to execute tests of arbitrary projects. ACM on Software Engineering, Vol. 2, pp. 1054-1076. Article

    Bui T. L., Dam H. K., & Hoda R. (2025, November). An LLM-based multi-agent framework for agile effort estimation. 40th IEEE/ACM International Conference on Automated Software Engineering, pp. 1032-1043. Article

    Chen X., Gao C., Chen C., Zhang G., & Liu Y. (2025). An empirical study on challenges for LLM application developers. ACM Transactions on Software Engineering and Methodology, Vol. 34, Issue 7, pp. 1-37. Article

    Dalpiaz F., & Niu N. (2020). Requirements engineering in the days of artificial intelligence. IEEE Software, Vol. 37, Issue 4, pp. 7-10. Article

    Daun M., & Brings J. (2023, June). How ChatGPT will change software engineering education. In Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education, Volume 1, pp. 110-116. Article

    Deng Y., Xia C. S., Yang C., Zhang S. D., Yang S., & Zhang L. (2024). Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries. 46th IEEE/ACM International Conference on Software Engineering, pp. 1-13. Article

    Dong Y., Jiang X., Jin Z., & Li G. (2024). Self-collaboration code generation via ChatGpt. ACM Transactions on Software Engineering and Methodology, Vol. 33, Issue 7, pp. 1-38. Article

    Doris A. C., Grandi D., Tomich R., Alam M. F., Ataei M., Cheong H., & Ahmed F. (2025). Designqa: A multimodal benchmark for evaluating large language models’ understanding of engineering documentation. Journal of Computing and Information Science in Engineering, Vol. 25, Issue 2, 021009. Article

    Elghariani K., & Kama N. (2016). Review on Agile requirements engineering challenges. 3rd International Conference on Computer and Information Sciences, pp. 507-512. Article

    Ezzini S., Abualhaija S., Arora C., & Sabetzadeh M. (2023). Ai-based question answering assistance for analyzing natural-language requirements. IEEE / ACM 45th International Conference on Software Engineering, pp. 1277-1289. Article

    Fan Z., Gao X., Mirchev M., Roychoudhury A., & Tan S. H. (2023). Automated repair of programs from large language models. IEEE/ACM 45th International Conference on Software Engineering, pp. 1469-1481. Article

    Fatima S., Ghaleb T. A., & Briand L. (2022). Flakify: A black-box, language model-based predictor for flaky tests. IEEE Transactions on Software Engineering, Vol. 49, Issue 4, pp. 1912-1927. Article

    Feldt R., Kang S., Yoon J., & Yoo S. (2023). Towards autonomous testing agents via conversational large language models. 38th IEEE/ACM International Conference on Automated Software Engineering, pp. 1688-1693. IEEE. Article

    Fu M., Tantithamthavorn C. K., Nguyen V., & Le T. (2023). ChatGpt for vulnerability detection, classification, and repair: How far are we? 30th Asia-Pacific Software Engineering Conference, pp. 632-636. Article

    Guo Y., Patsakis C., Hu Q., Tang Q., & Casino F. (2024). Outside the comfort zone: Analysing llm capabilities in software vulnerability detection. European Symposium on Research in Computer Security, pp. 271-289. Cham: Springer Nature Switzerland. Article

    Han H., Kim J., Yoo J., Lee Y., & Hwang S. W. (2024). Archcode: Incorporating software requirements in code generation with large language models. Annual Meeting of the Association for Computational Linguistics, Vol. 1, pp. 13520-13552. Article

    Heyn H. M., Knauss E., Muhammad A. P., Eriksson O., Linder J., Subbiah P., ... & Tungal S. (2021, May). Requirement engineering challenges for ai-intense systems development. IEEE/ACM 1st Workshop on AI Engineering-Software Engineering for AI, pp. 89-96. IEEE. Article

    Jain N., Vaidyanath S., Iyer A., Natarajan N., Parthasarathy S., Rajamani S., & Sharma R. (2022). Jigsaw: Large language models meet program synthesis. 44th International Conference on Software Engineering, pp. 1219-1231. Article

    Jain N., Vaidyanath S., Iyer A., Natarajan N., Parthasarathy S., Rajamani S., & Sharma R. (2022). Jigsaw: Large language models meet program synthesis. 44th International Conference on Software Engineering, pp. 1219-1231. Article

    Jimenez C. E., Yang J., Wettig A., Yao S., Pei K., Press O., & Narasimhan K. (2024, May). Swe-bench: Can language models resolve real-world github issues? International Conference on Learning Representations, Vol. 2024, pp. 54107-54157.

    Jin M., Shahriar S., Tufano M., Shi X., Lu S., Sundaresan N., & Svyatkovskiy A. (2023). Inferfix: End-to-end program repair with LLMs. 31st ACM Joint European Software Engineering Conference and Symposium on The Foundations of Software Engineering, pp. 1646-1656. Article

    Kaddour J., Harris J., Mozes M., Bradley H., Raileanu R., & McHardy R. (2023). Challenges and applications of large language models. arXiv preprint arXiv:2307.10169. https://doi.org/10.48550/arXiv.2307.10169 Article

    Kang S., Yoon J., & Yoo S. (2023, May). Large language models are few-shot testers: Exploring LLM-based general bug reproduction. IEEE/ACM 45th International Conference on Software Engineering, pp. 2312-2323. Article

    Khare A., Dutta S., Li Z., Solko-Breslin A., Alur R., & Naik M. (2025). Understanding the effectiveness of large language models in detecting security vulnerabilities. IEEE Conference on Software Testing, Verification and Validation, pp. 103-114. Article

    Khattab O., Singhvi A., Maheshwari P., Zhang Z., Santhanam K., Vardhamanan S., ... & Potts C. (2023). Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714. Article

    Khojah R., Mohamad M., Leitner P., & de Oliveira Neto F. G. (2024). Beyond code generation: An observational study of chatgpt usage in software engineering practice. ACM on Software Engineering, Vol. 1(FSE), pp. 1819-1840. Article

    Kozov V., Ivanova G., & Atanasova D. (2024). Practical Application of AI and Large Language Models in Software Engineering Education. International Journal of Advanced Computer Science & Applications, Vol. 15, Issue 1. Article

    Krishna M., Gaur B., Verma A., & Jalote P. (2024, June). Using LLMs in software requirements specifications: An empirical evaluation. 32nd International Requirements Engineering Conference, pp. 475-483. Article

    Li J., Li G., Li Y., & Jin Z. (2025). Structured chain-of-thought prompting for code generation. ACM Transactions on Software Engineering and Methodology, Vol. 34, Issue 2, pp. 1-23. Article

    Li J., Rabbi F., Cheng C., Sangalay A., Tian Y., & Yang J. (2026). An exploratory study on fine-tuning large language models for secure code generation. Empirical Software Engineering, Vol. 31, Issue 4. Article

    Li J., Zhao Y., Li Y., Li G., & Jin Z. (2024). Acecoder: An effective prompting technique specialized in code generation. ACM Transactions on Software Engineering and Methodology, Vol. 33, Issue 8, pp. 1-26. Article

    Liu J., Xia C. S., Wang Y., & Zhang L. (2023). Is your code generated by ChatGpt really, correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems, Vol. 36, pp. 21558-21572. Article

    Lops A., Narducci F., Ragone A., Trizio M., & Bartolini C. (2025, March). A system for automated unit test generation using large language models and assessment of generated test suites. IEEE International Conference on Software Testing, Verification and Validation Workshops, pp. 29-36. Article

    Lu Q., Zhu L., Xu X., Xing Z., Harrer S., & Whittle J. (2024). Towards responsible generative ai: A reference architecture for designing foundation model-based agents. IEEE 21st International Conference on Software Architecture Companion, pp. 119-126. Article

    Luitel D., Hassani S., & Sabetzadeh M. (2023). Using language models for enhancing the completeness of natural-language requirements. In International Working Conference on Requirements Engineering: Foundation for Software Quality, pp. 87-104. Cham: Springer Nature Switzerland. Article

    Manish S. (2024). An autonomous multi-agent LLM framework for agile software development. International Journal of Trend in Scientific Research and Development, Vol. 8, Issue 5, pp. 892-898.Article

    Mathews N. S., & Nagappan M. (2024, October). Test-driven development and llm-based code generation. 39th IEEE/ACM International Conference on Automated Software Engineering, pp. 1583-1594. Article

    Minaee S., Mikolov T., Nikzad N., Chenaghlu M., Socher R., Amatriain X., & Gao J. (2024). Large language models: A survey. arXiv preprint arXiv:2402.06196. Article

    Molina F., Gorla A., & d’Amorim M. (2025). Test Oracle Automation in the era of LLMs. ACM Transactions on Software Engineering and Methodology, Vol. 34, Issue 5, pp. 1-24. Article

    Mu F., Shi L., Wang S., Yu Z., Zhang B., Wang C., ... & Wang Q. (2024). Clarifygpt: A framework for enhancing llm-based code generation via requirements clarification. ACM on Software Engineering, 1, pp. 2332-2354. Article

    Nam D., Macvean A., Hellendoorn V., Vasilescu B., & Myers B. (2024, April). Using an LLM to help with code understanding. IEEE/ACM 46th International Conference on Software Engineering, pp. 1-13. Article

    Neelapu M. Enhancing Software Testing Efficiency with Generative AI and Large Language Models. IJLRP-International Journal of Leading Research Publication, Vol. 5, Issue 12. Article

    Noever D. (2023). Can large language models find and fix vulnerable software? arXiv preprint arXiv:2308.10345. Article

    Nong Y., Aldeen M., Cheng L., Hu H., Chen F., & Cai H. (2024). Chain-of-thought prompting of large language models for discovering and fixing software vulnerabilities. arXiv preprint arXiv:2402.17230. Article

    Ozkaya I. (2023). Application of large language models to software engineering tasks: Opportunities, risks, and implications. IEEE software, Vol. 40, Issue 3, pp. 4-8. Article

    Page M. J., McKenzie J. E., Bossuyt P. M., Boutron I., Hoffmann T. C., Mulrow C. D., ... & Moher D. (2021). The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. bmj, Vol. 372. Article

    Plein L., Ouédraogo W. C., Klein J., & Bissyandé T. F. (2024). Automatic generation of test cases based on bug reports: a feasibility study with large language models. IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, pp. 360-361. Article

    Pudari R., & Ernst N. A. (2023). From copilot to pilot: Towards AI supported software development. arXiv preprint arXiv:2303.04142. Article

    Rajbhoj A., Somase A., Kulkarni P., & Kulkarni V. (2024). Accelerating software development using generative AI: ChatGPT case study. 17th Innovations in Software Engineering Conference, pp. 1-11. Article

    Rao N., Jain K., Alon U., Le Goues C., & Hellendoorn V. J. (2023). CAT-LM training language models on aligned code and tests. 38th IEEE/ACM International Conference on Automated Software Engineering, pp. 409-420. Article

    Rodriguez A. D., Dearstyne K. R., & Cleland-Huang J. (2023). Prompts matter: Insights and strategies for prompt engineering in automated software traceability. 31st International Requirements Engineering Conference Workshops, pp. 455-464. Article

    Rodriguez-Cardenas D. (2025, June). Towards More Interpretable Large Language Models for Code. 33rd ACM International Conference on the Foundations of Software Engineering, pp. 1270-1272. Article

    Ronanki K., Berger C., & Horkoff J. (2023). Investigating chatgpt’s potential to assist in requirements elicitation processes. 49th IEEE Euromicro Conference on Software Engineering and Advanced Applications, pp. 354-361. Article

    Ronanki K., Cabrero-Daniel B., Horkoff J., & Berger C. (2024). Requirements engineering using generative ai: Prompts and prompting patterns. Generative AI for Effective Software Development, pp. 109-127. Cham: Springer Nature Switzerland. Article

    Ross S. I., Martinez F., Houde S., Muller M., & Weisz J. D. (2023, March). The programmer’s assistant: Conversational interaction with a large language model for software development. 28th International Conference on Intelligent User Interfaces, pp. 491-514. Article

    Sallou J., Durieux T., & Panichella A. (2024, April). Breaking the silence: the threats of usingLLMs in software engineering. ACM/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results, pp. 102-106. Article

    Santos R., Santos I., Magalhaes C., & de Souza Santos R. (2024). Are we testing or being tested? exploring the practical applications of large language models in software testing. IEEE Conference on Software Testing, Verification and Validation, pp. 353-360. Article

    Shin J., Tang C., Mohati T., Nayebi M., Wang S., & Hemmati H. (2023). Prompt engineering or fine tuning: An empirical assessment of large language models in automated software engineering tasks. arXiv preprint arXiv:2310.10508. Article

    Shirafuji A., Oda Y., Suzuki J., Morishita M., & Watanobe Y. (2023, December). Refactoring programs using large language models with few-shot examples. 30th Asia-Pacific Software Engineering Conference, pp. 151-160. Article

    Siddiq M. L., Da Silva Santos J. C., Tanvir R. H., Ulfat N., Al Rifat F., & Carvalho Lopes V. (2024, June). Using large language models to generate junit tests: An empirical study. 28th International Conference on Evaluation and Assessment in Software Engineering, pp. 313-322. Article

    Singla T., Anandayuvaraj D., Kalu K. G., Schorlemmer T. R., & Davis J. C. (2023). An empirical study on using large language models to analyze software supply chain security failures. Workshop on Software Supply Chain Offensive Research and Ecosystem Defenses, pp. 5-15. Article

    Son H. S. (2025). Using large language models in software requirements analysis. Edelweiss Appl. Sci. Technol, Vol. 9, Issue 3, pp. 856-863.

    Subahi, A. F. (2026). Prompt Engineering for Advancing Software Engineering: Using Generative Language Models in the Development of Domain-Specific Languages. IEEE Access, Vol. 14, pp. 22834-22850. Article

    Sun W., Miao Y., Li Y., Zhang H., Fang C., Liu Y., ... & Chen Z. (2025). Source code summarization in the era of large language models. IEEE/ACM 47th International Conference on Software Engineering, pp. 1882-1894. Article

    Tihanyi N., Bisztray T., Jain R., Ferrag M. A., Cordeiro L. C., & Mavroeidis V. (2023). The formai dataset: Generative ai in software security through the lens of formal verification. 19th International Conference on Predictive Models and Data Analytics in Software Engineering, pp. 33-43. Article

    Tihanyi N., Charalambous Y., Jain R., Ferrag M. A., & Cordeiro L. C. (2025). A new era in software security: Towards self-healing software via large language models and formal verification. IEEE/ACM International Conference on Automation of Software Test, pp. 136-147. IEEE. Article

    Treude C., & Hata H. (2023). She elicits requirements and he tests: Software engineering gender bias in large language models. IEEE/ACM 20th International Conference on Mining Software Repositories, pp. 624-629. Article

    Wang Y., Kordi Y., Mishra S., Liu A., Smith N. A., Khashabi D., & Hajishirzi H. (2023). Self-instruct: Aligning language models with self-generated instructions. 61st Annual Meeting of the Association for Computational Linguistics, Vol. 1, pp. 13484-13508. Article

    Wang Z., Liu K., Li G., & Jin Z. (2024). Hits: High-coverage llm-based unit test generation via method slicing. 39th IEEE/ACM International Conference on Automated Software Engineering, pp. 1258-1268. Article

    Wei B. (2024). Requirements are all you need: From requirements to code with LLM. In 2024 IEEE 32nd International Requirements Engineering Conference, pp. 416-422. Article

    White J., Hays S., Fu Q., Spencer-Smith J., & Schmidt D. C. (2024). Chatgpt prompt patterns for improving code quality, refactoring, requirements elicitation, and software design. In Generative ai for effective software development, pp. 71-108. Cham: Springer Nature Switzerland. Article

    Xia C. S., Paltenghi M., Le Tian J., Pradel M., & Zhang L. (2024). Fuzz4all: Universal fuzzing with large language models. IEEE/ACM 46th International Conference on Software Engineering, pp. 1-13. Article

    Xue Z., Li L., Tian S., Chen X., Li P., Chen L., ... & Zhang M. (2024). Llm4fin: Fully automating LLM-powered test case generation for fintech software acceptance testing. 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, pp. 1643-1655. Article

    Yang A. Z., Le Goues C., Martins R., & Hellendoorn V. (2024). Large language models for test-free fault localization. 46th IEEE/ACM International Conference on Software Engineering, pp. 1-12. Article

    Yang Z., Liu F., Yu Z., Keung J. W., Li J., Liu S., ... & Li G. (2024). Exploring and unleashing the power of large language models in automated code translation. ACM on Software Engineering, Vol. 1(FSE), pp. 1585-1608. Article

    Yu S., Fang C., Ling Y., Wu C., & Chen Z. (2023). LLM for test script generation and migration: Challenges, capabilities, and opportunities. IEEE 23rd International Conference on Software Quality, Reliability, and Security, pp. 206-217. Article

    Zhang Y., Ruan H., Fan Z., & Roychoudhury A. (2024, September). Autocoderover: Autonomous program improvement. 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, pp. 1592-1604. Article

    Zhang Y., Xie Y., Lit S., Liu K., Wang C., Jia Z., ... & Liao X. (2025, April). Unseen horizons: Unveiling the real capability of LLM code generation beyond the familiar. IEEE/ACM 47th International Conference on Software Engineering, pp. 604-615. Article

    Zhang Z., Rayhan M., Herda T., Goisauf M., & Abrahamsson P. (2024). Llm-based agents for automating the enhancement of user story quality: An early report. International Conference on Agile Software Development, pp. 117-126. Cham: Springer Nature Switzerland. Article

    Zheng Z., Ning K., Zhong Q., Chen J., Chen W., Guo L., ... & Wang Y. (2025). Towards an understanding of large language models in software engineering tasks. Empirical Software Engineering, Vol. 30, Issue 2. Article

    Zhou X., Zhang T., & Lo D. (2024). Large language model for vulnerability detection: Emerging results and future directions. ACM/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results, pp. 47-51. Article

    Zhu K., Wang J., Zhou J., Wang Z., Chen H., Wang Y., & Xie X. (2023). Promptrobust: Towards evaluating the robustness of large language models on adversarial prompts. ACM Workshop on Large AI Systems and Models with Privacy and Safety Analysis, pp. 57-68. Article

    Ziegler D. M., Stiennon N., Wu J., Brown T. B., Radford A., Amodei D., ... & Irving G. (2019). Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593. Article