Evaluating Retrieval Quality in Retrieval-Augmented Generation
Оценка качества извлечения в системах генерации с дополненным извлечением
2024-07-10
SCID: 54.1/x67vtw5f
Discuss with AI
Kendall's tau correlationdocument-level relevanceeRAGretrieval quality evaluationretrieval-augmented generation
Figures from the paper
Abstract (AI)
Evaluating retrieval-augmented generation (RAG) presents challenges, particularly for retrieval models within these systems. Traditional end-to-end evaluation methods are computationally expensive. Furthermore, evaluation of the retrieval model's performance based on query-document relevance labels shows a small correlation with the RAG system's downstream performance. We propose a novel evaluation approach, eRAG, where each document in the retrieval list is individually utilized by the large language model within the RAG system. The output generated for each document is then evaluated based on the downstream task ground truth labels. In this manner, the downstream performance for each document serves as its relevance label. We employ various downstream task metrics to obtain document-level annotations and aggregate them using set-based or ranking metrics. Extensive experiments on a wide range of datasets demonstrate that eRAG achieves a higher correlation with downstream RAG performance compared to baseline methods, with improvements in Kendall's tau correlation ranging from 0.168 to 0.494. Additionally, eRAG offers significant computational advantages, improving runtime and consuming up to 50 times less GPU memory than end-to-end evaluation.
Key Findings
1
Across diverse datasets, eRAG correlates more strongly with downstream RAG performance than baseline evaluation methods, improving Kendall’s tau by 0.168–0.494.
2
The method derives document-level relevance annotations from downstream ground-truth metrics rather than conventional query-document relevance labels.
3
eRAG evaluates retrieval quality by feeding each retrieved document individually to the language model and labeling it using downstream task performance.
4
eRAG provides substantial computational benefits over end-to-end evaluation, including up to 50× lower GPU memory consumption and improved runtime.
Research Object
retrieval models within retrieval-augmented generation (RAG) systems
Research Subject
the correlation between document-level retrieval relevance evaluations and downstream RAG performance, including evaluation efficiency
Publication Details
Publication Date
2024-07-10
Journal
Publisher
ISSN
Cited by
146
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest