CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
CRUD-RAG: комплексный китайский бенчмарк для генерации с дополнением извлечённой информацией большими языковыми моделями
2024-10-19
SCID: 54.1/4hur3d7g
Discuss with AI
CRUD-RAGChinese benchmarkRAG evaluationlarge language modelsretrieval-augmented generation
Figures from the paper
Abstract (AI)
Retrieval-augmented generation (RAG) is a technique that enhances the capabilities of large language models (LLMs) by incorporating external knowledge sources. This method addresses common LLM limitations, including outdated information and the tendency to produce inaccurate “hallucinated” content. However, evaluating RAG systems is a challenge. Most benchmarks focus primarily on question-answering applications, neglecting other potential scenarios where RAG could be beneficial. Accordingly, in the experiments, these benchmarks often assess only the LLM components of the RAG pipeline or the retriever in knowledge-intensive scenarios, overlooking the impact of external knowledge base construction and the retrieval component on the entire RAG pipeline in non-knowledge-intensive scenarios. To address these issues, this article constructs a large-scale and more comprehensive benchmark and evaluates all the components of RAG systems in various RAG application scenarios. Specifically, we refer to the CRUD actions that describe interactions between users and knowledge bases and also categorize the range of RAG applications into four distinct types—create, read, update, and delete (CRUD). “Create” refers to scenarios requiring the generation of original, varied content. “Read” involves responding to intricate questions in knowledge-intensive situations. “Update” focuses on revising and rectifying inaccuracies or inconsistencies in pre-existing texts. “Delete” pertains to the task of summarizing extensive texts into more concise forms. For each of these CRUD categories, we have developed different datasets to evaluate the performance of RAG systems. We also analyze the effects of various components of the RAG system, such as the retriever, context length, knowledge base construction, and LLM. Finally, we provide useful insights for optimizing the RAG technology for different scenarios. The source code is available at GitHub: https://github.com/IAAR-Shanghai/CRUD_RAG .
Key Findings
1
CRUD-RAG introduces a large-scale Chinese benchmark covering four RAG application categories: create, read, update, and delete.
2
Experiments analyze how retriever choice, context length, knowledge-base construction, and LLM selection affect RAG performance.
3
Separate datasets are developed to assess original content generation, knowledge-intensive question answering, text correction, and long-text summarization.
4
The benchmark evaluates complete RAG pipelines across diverse scenarios, rather than focusing only on question answering, retrievers, or LLM components.
5
The study reports scenario-specific insights for optimizing RAG systems and highlights the importance of evaluating external knowledge construction and retrieval beyond knowledge-intensive tasks.
Research Object
CRUD-RAG benchmark and the retrieval-augmented generation (RAG) systems evaluated across create, read, update, and delete application scenarios
Research Subject
Comprehensive evaluation of the end-to-end RAG pipeline, including retriever performance, context length, knowledge-base construction, and LLM effects across diverse application scenarios
Publication Details
Publication Date
2024-10-19
Journal
Publisher
ISSN
Cited by
95
Access Type
Author Information
Download PDF
Subscribe to digest