A Survey on Large Language Models for Code Generation
Обзор больших языковых моделей для генерации кода
2025-07-08
SCID: 54.1/n5vx4u9r
Discuss with AI
Code LLMsHumanEval benchmarkMBPP benchmarkcode generationlarge language models
Figures from the paper
Abstract (AI)
Large Language Models (LLMs) have garnered remarkable advancements across diverse code-related tasks, known as Code LLMs, particularly in code generation that generates source code with LLM from natural language descriptions. This burgeoning field has captured significant interest from both academic researchers and industry professionals due to its practical significance in software development, e.g., GitHub Copilot . Despite the active exploration of LLMs for a variety of code tasks, either from the perspective of Natural Language Processing (NLP) or Software Engineering (SE) or both, there is a noticeable absence of a comprehensive and up-to-date literature review dedicated to LLM for code generation. In this survey, we aim to bridge this gap by providing a systematic literature review that serves as a valuable reference for researchers investigating the cutting-edge progress in LLMs for code generation. We introduce a taxonomy to categorize and discuss the recent developments in LLMs for code generation, covering aspects such as data curation, latest advances, performance evaluation, ethical implications, environmental impact, and real-world applications. In addition, we present a historical overview of the evolution of LLMs for code generation and provide a quantitative and qualitative comparative analysis of experimental results of code LLMs, sourced from their original papers to ensure a fair comparison on the HumanEval, MBPP, and BigCodeBench benchmarks, across various levels of difficulty and types of programming tasks, to highlight the progressive enhancements in LLM capabilities for code generation. We identify critical challenges and promising opportunities regarding the gap between academia and practical development. Furthermore, we have established a dedicated resource GitHub page ( https://github.com/juyongjiang/CodeLLMSurvey ) to continuously document and disseminate the most recent advances in the field.
Key Findings
1
Comparisons span different difficulty levels and programming-task types, highlighting progressive improvements in LLM capabilities for code generation.
2
It introduces a taxonomy covering data curation, technical advances, performance evaluation, ethical implications, environmental impact, and real-world applications of code LLMs.
3
The study provides a historical account of code-generation LLM evolution and compares experimental results quantitatively and qualitatively using HumanEval, MBPP, and BigCodeBench.
4
The survey addresses the lack of a comprehensive, up-to-date literature review focused specifically on large language models for code generation.
5
The survey identifies persistent challenges and promising opportunities, particularly concerning the gap between academic research and practical software development, and maintains a continuously updated GitHub resource.
Research Object
large language models for code generation
Research Subject
their development, capabilities, evaluation, applications, ethical and environmental implications, and remaining challenges
Publication Details
Publication Date
2025-07-08
Journal
Publisher
ISSN
Cited by
205
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai8
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing2022
A Survey on Evaluation of Large Language Models2024
CodeBERT: A Pre-Trained Model for Programming and Natural Languages2020
A survey on large language model based autonomous agents2024
Software Testing With Large Language Models: Survey, Landscape, and Vision2024
Large Language Models for Software Engineering: Survey and Open Problems2023
GPT-NeoX-20B: An Open-Source Autoregressive Language Model2022
Benchmarking Large Language Models in Retrieval-Augmented Generation2024