LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation

Галлюцинации больших языковых моделей при практической генерации кода: проявления, механизм и методы устранения
Ziyao Zhang, Chong Wang, Yanlin Wang, Ensheng Shi, Yuchi Ma, Wanjun Zhong, Jiachi Chen, Mingzhi Mao, Zibin Zheng
2025-06-22

LLM hallucinationscode generationhallucination taxonomyrepository-level generationretrieval-augmented generation
Code generation aims to automatically generate code from input requirements, significantly enhancing development efficiency. Recent large language models (LLMs) based approaches have shown promising results and revolutionized code generation task. Despite the promising performance, LLMs often generate contents with hallucinations, especially for the code generation scenario requiring the handling of complex contextual dependencies in practical development process. Although previous study has analyzed hallucinations in LLM-powered code generation, the study is limited to standalone function generation. In this paper, we conduct an empirical study to study the phenomena, mechanism, and mitigation of LLM hallucinations within more practical and complex development contexts in repository-level generation scenario. First, we manually examine the code generation results from six mainstream LLMs to establish a hallucination taxonomy of LLM-generated code. Next, we elaborate on the phenomenon of hallucinations, analyze their distribution across different models. We then analyze causes of hallucinations and identify four potential factors contributing to hallucinations. Finally, we propose an RAG-based mitigation method, which demonstrates consistent effectiveness in all studied LLMs.
1
An RAG-based mitigation method consistently reduces hallucinations across all studied LLMs.
2
Manual analysis of outputs from six mainstream LLMs establishes a taxonomy of hallucinations in generated code.
3
The authors characterize hallucination phenomena and compare their distribution across different LLMs.
4
The study examines LLM hallucinations in repository-level code generation, extending prior work beyond standalone function generation.
5
The study identifies four potential factors contributing to hallucinations in practical code generation contexts.

LLM-generated code in repository-level code generation

Hallucination phenomena, mechanisms, contributing factors, and mitigation effectiveness

Publication Details
Publication Date
2025-06-22
Journal
Publisher
ISSN
Cited by
91
Access Type
Author Information
Authors
Ziyao Zhang
Chong Wang
Yanlin Wang
Ensheng Shi
Yuchi Ma
Wanjun Zhong
Jiachi Chen
Mingzhi Mao
Zibin Zheng
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%