Benchmarking Large Language Models in Retrieval-Augmented Generation

Сравнительный анализ больших языковых моделей в генерации с дополнением извлечённой информацией
Hongyu Lin, Jiawei Chen, Xianpei Han, Le Sun
2024-03-24

Retrieval-Augmented GenerationRetrieval-Augmented Generation Benchmarkcounterfactual robustnesslarge language modelsnoise robustness
Retrieval-Augmented Generation (RAG) is a promising approach for mitigating the hallucination of large language models (LLMs). However, existing research lacks rigorous evaluation of the impact of retrieval-augmented generation on different large language models, which make it challenging to identify the potential bottlenecks in the capabilities of RAG for different LLMs. In this paper, we systematically investigate the impact of Retrieval-Augmented Generation on large language models. We analyze the performance of different large language models in 4 fundamental abilities required for RAG, including noise robustness, negative rejection, information integration, and counterfactual robustness. To this end, we establish Retrieval-Augmented Generation Benchmark (RGB), a new corpus for RAG evaluation in both English and Chinese. RGB divides the instances within the benchmark into 4 separate testbeds based on the aforementioned fundamental abilities required to resolve the case. Then we evaluate 6 representative LLMs on RGB to diagnose the challenges of current LLMs when applying RAG. Evaluation reveals that while LLMs exhibit a certain degree of noise robustness, they still struggle significantly in terms of negative rejection, information integration, and dealing with false information. The aforementioned assessment outcomes indicate that there is still a considerable journey ahead to effectively apply RAG to LLMs.
1
Large language models demonstrate some robustness to retrieval noise but struggle substantially with negative rejection, information integration, and false information.
2
RGB evaluates four fundamental RAG abilities: noise robustness, negative rejection, information integration, and counterfactual robustness.
3
The findings indicate that current large language models remain far from reliably supporting effective RAG applications.
4
The paper introduces RGB, a bilingual English–Chinese benchmark for systematically evaluating retrieval-augmented generation (RAG).
5
The study benchmarks six representative large language models to diagnose model-specific challenges in applying RAG.

Large language models using Retrieval-Augmented Generation (RAG)

their performance in noise robustness, negative rejection, information integration, and counterfactual robustness

Publication Details
Publication Date
2024-03-24
Journal
Publisher
ISSN
Cited by
380
Access Type
Author Information
Authors
Hongyu Lin
Jiawei Chen
Xianpei Han
Le Sun
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%