Large Language Models for Software Engineering: A Systematic Literature Review

Большие языковые модели для программной инженерии: систематический обзор литературы
Xiapu Luo, John Grundy, David Lo, Haoyu Wang, Yue Liu, Kailong Wang, Li Li, Yanjie Zhao, Xinyi Hou, Yang Zhou
2023-08-21

LLM4SElarge language modelssoftware engineeringsoftware engineering taskssystematic literature review
Large Language Models (LLMs) have significantly impacted numerous domains, including Software Engineering (SE). Many recent publications have explored LLMs applied to various SE tasks. Nevertheless, a comprehensive understanding of the application, effects, and possible limitations of LLMs on SE is still in its early stages. To bridge this gap, we conducted a systematic literature review (SLR) on LLM4SE, with a particular focus on understanding how LLMs can be exploited to optimize processes and outcomes. We select and analyze 395 research papers from January 2017 to January 2024 to answer four key research questions (RQs). In RQ1, we categorize different LLMs that have been employed in SE tasks, characterizing their distinctive features and uses. In RQ2, we analyze the methods used in data collection, preprocessing, and application, highlighting the role of well-curated datasets for successful LLM for SE implementation. RQ3 investigates the strategies employed to optimize and evaluate the performance of LLMs in SE. Finally, RQ4 examines the specific SE tasks where LLMs have shown success to date, illustrating their practical contributions to the field. From the answers to these RQs, we discuss the current state-of-the-art and trends, identifying gaps in existing research, and flagging promising areas for future study. Our artifacts are publicly available at https://github.com/xinyi-hou/LLM4SE_SLR.
1
A systematic literature review analyzed 395 studies on large language models for software engineering published from January 2017 to January 2024.
2
Successful LLM4SE implementation depends substantially on data collection, preprocessing, and the availability of well-curated datasets.
3
The review categorizes LLMs used in software engineering, including their distinguishing characteristics and applications across tasks.
4
The review identifies software engineering tasks where LLMs have demonstrated practical success, while highlighting research gaps and promising future directions.
5
The study synthesizes strategies for optimizing and evaluating LLM performance in software engineering tasks.

Application of large language models (LLMs) to software engineering tasks

The applications, effects, optimization and evaluation strategies, successes, and limitations of Large Language Models in Software Engineering

Publication Details
Publication Date
2023-08-21
Journal
Publisher
ISSN
Cited by
119
Access Type
Author Information
Authors
Xiapu Luo
John Grundy
David Lo
Haoyu Wang
Yue Liu
Kailong Wang
Li Li
Yanjie Zhao
Xinyi Hou
Yang Zhou
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%