Bias in Large Language Models: Origin, Evaluation, and Mitigation
Предвзятость в больших языковых моделях: происхождение, оценка и методы снижения
2026-04-24
SCID: 54.1/6b5hrpte
Discuss with AI
LLM biasbias evaluationbias mitigationfair artificial intelligencelarge language models
Figures from the paper
Abstract (AI)
Large language models (LLMs) have revolutionized natural language processing, but their susceptibility to biases poses significant challenges. This comprehensive review examines the landscape of bias in LLMs, from its origins to current mitigation strategies. We categorize biases as intrinsic and extrinsic, analyzing their manifestations in various natural language processing (NLP) tasks. The review critically assesses a range of bias evaluation methods, including data-level, model-level, and output-level approaches, providing researchers with a robust toolkit for bias detection. We further explore mitigation strategies, categorizing them into pre-model, intra-model, and post-model techniques, highlighting their effectiveness and limitations. Ethical and legal implications of biased LLMs are discussed, emphasizing potential harms in real-world applications such as healthcare and criminal justice. By synthesizing current knowledge on bias in LLMs, this review contributes to the ongoing effort to develop fair and responsible artificial intelligence (AI) systems. Our work serves as a comprehensive resource for researchers and practitioners working towards understanding, evaluating, and mitigating bias in LLMs, fostering the development of more equitable AI technologies.
Key Findings
1
Biased LLMs can cause significant ethical and legal harms in real-world applications, including healthcare and criminal justice.
2
It evaluates bias-detection methods at the data, model, and output levels, providing a structured toolkit for assessing LLM bias.
3
Mitigation strategies are organized into pre-model, intra-model, and post-model techniques, with their effectiveness and limitations critically examined.
4
The review categorizes LLM bias into intrinsic and extrinsic forms and examines their manifestations across diverse NLP tasks.
5
The review synthesizes current knowledge to support the development of fairer and more responsible AI systems.
Research Object
bias in large language models (LLMs)
Research Subject
the origins, manifestations, evaluation, and mitigation of bias in LLMs, including ethical and legal implications
Publication Details
Publication Date
2026-04-24
Journal
Publisher
ISSN
Cited by
26
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest