Benchmarking Large Language Models in Retrieval-Augmented Generation
Сравнительный анализ больших языковых моделей в генерации с дополнением извлечённой информацией
2024-03-24
SCID: 54.1/epde4uqw
Discuss with AI
Retrieval-Augmented GenerationRetrieval-Augmented Generation Benchmarkcounterfactual robustnesslarge language modelsnoise robustness
Figures from the paper
Abstract (AI)
Retrieval-Augmented Generation (RAG) is a promising approach for mitigating the hallucination of large language models (LLMs). However, existing research lacks rigorous evaluation of the impact of retrieval-augmented generation on different large language models, which make it challenging to identify the potential bottlenecks in the capabilities of RAG for different LLMs. In this paper, we systematically investigate the impact of Retrieval-Augmented Generation on large language models. We analyze the performance of different large language models in 4 fundamental abilities required for RAG, including noise robustness, negative rejection, information integration, and counterfactual robustness. To this end, we establish Retrieval-Augmented Generation Benchmark (RGB), a new corpus for RAG evaluation in both English and Chinese. RGB divides the instances within the benchmark into 4 separate testbeds based on the aforementioned fundamental abilities required to resolve the case. Then we evaluate 6 representative LLMs on RGB to diagnose the challenges of current LLMs when applying RAG. Evaluation reveals that while LLMs exhibit a certain degree of noise robustness, they still struggle significantly in terms of negative rejection, information integration, and dealing with false information. The aforementioned assessment outcomes indicate that there is still a considerable journey ahead to effectively apply RAG to LLMs.
Key Findings
1
Large language models demonstrate some robustness to retrieval noise but struggle substantially with negative rejection, information integration, and false information.
2
RGB evaluates four fundamental RAG abilities: noise robustness, negative rejection, information integration, and counterfactual robustness.
3
The findings indicate that current large language models remain far from reliably supporting effective RAG applications.
4
The paper introduces RGB, a bilingual English–Chinese benchmark for systematically evaluating retrieval-augmented generation (RAG).
5
The study benchmarks six representative large language models to diagnose model-specific challenges in applying RAG.
Research Object
Large language models using Retrieval-Augmented Generation (RAG)
Research Subject
their performance in noise robustness, negative rejection, information integration, and counterfactual robustness
Publication Details
Publication Date
2024-03-24
Journal
Publisher
ISSN
Cited by
380
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai6
Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Meta LLM-Integrated Systems2020
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering2021
REALM: Retrieval-Augmented Language Model Pre-Training2020
Atlas: Few-shot Learning with Retrieval Augmented Language Models2022
Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture2026
Retrieval-guided Dialogue Response Generation via a Matching-to-Generation Framework2019
Cited by7
Generative AI Agents With Large Language Model for Satellite Networks via a Mixture of Experts Transmission2024
A Survey on Large Language Models for Code Generation2025
Retrieval-Augmented Generation (RAG) Chatbots for Education: A Survey of Applications2025
Generative AI and the Future of Democratic Citizenship2024
Generative AI for Facial Expressions in 3D Game Characters: A Retrieval-Augmented Approach2025
ChronoGrapher: Event-Centric Knowledge Graph Construction via Informed Graph Traversal2025
SpaRAGraph: Spatial Reasoning using Retrieval-Augmented Generation2026