Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Песнь сирен в океане ИИ: обзор галлюцинаций в больших языковых моделях
Yue Zhang, Longyue Wang, Shuming Shi, Yafu Li, Leyang Cui, Wei Bi, Xinting Huang, Deng Cai, Lemao Liu, Tingchen Fu, Enbo Zhao, Yu Zhang, Xu, Chen, Yulong Chen, Anh Tuan Luu, Freda Shi
2023-09-03

evaluation benchmarks for hallucinationhallucination detectionhallucination explanationhallucination in large language modelshallucination mitigation
While large language models (LLMs) have demonstrated remarkable capabilities across a range of downstream tasks, a significant concern revolves around their propensity to exhibit hallucinations: LLMs occasionally generate content that diverges from the user input, contradicts previously generated context, or misaligns with established world knowledge. This phenomenon poses a substantial challenge to the reliability of LLMs in real-world scenarios. In this paper, we survey recent efforts on the detection, explanation, and mitigation of hallucination, with an emphasis on the unique challenges posed by LLMs. We present taxonomies of the LLM hallucination phenomena and evaluation benchmarks, analyze existing approaches aiming at mitigating LLM hallucination, and discuss potential directions for future research.
1
Hallucination in LLMs poses a major challenge to their reliability in real-world applications.
2
LLMs frequently exhibit hallucinations: generating content that diverges from user input, contradicts prior context, or misaligns with world knowledge.
3
The authors identify open directions for future research on detecting, explaining, and mitigating hallucination in LLMs.
4
The paper provides taxonomies of hallucination phenomena and of evaluation benchmarks tailored to LLM-specific challenges.
5
The survey analyzes existing detection, explanation, and mitigation approaches for LLM hallucination.

Hallucination phenomena in large language models (LLMs)

Detection, explanation, evaluation, and mitigation of LLM hallucinations including taxonomies, benchmarks, analysis of approaches, and research directions

Publication Details
Publication Date
2023-09-03
Journal
Publisher
ISSN
Cited by
244
Access Type
Author Information
Authors
Yue Zhang
Longyue Wang
Shuming Shi
Yafu Li
Leyang Cui
Wei Bi
Xinting Huang
Deng Cai
Lemao Liu
Tingchen Fu
Enbo Zhao
Yu Zhang
Xu, Chen
Yulong Chen
Anh Tuan Luu
Freda Shi
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%