Benchmarking Retrieval-Augmented Generation for Medicine
Бенчмаркинг генерации с дополнением извлечённой информацией в медицине
2024-01-01
SCID: 54.1/7474bj5w
Discuss with AI
MEDRAG toolkitMIRAGE benchmarklost-in-the-middle effectmedical question answeringmedical retrieval-augmented generation
Figures from the paper
Abstract (AI)
While large language models (LLMs) have achieved state-of-the-art performance on a wide range of medical question answering (QA) tasks, they still face challenges with hallucinations and outdated knowledge.Retrievalaugmented generation (RAG) is a promising solution and has been widely adopted.However, a RAG system can involve multiple flexible components, and there is a lack of best practices regarding the optimal RAG setting for various medical purposes.To systematically evaluate such systems, we propose the Medical Information Retrieval-Augmented Generation Evaluation (MIRAGE), a first-of-its-kind benchmark including 7,663 questions from five medical QA datasets.Using MIRAGE, we conducted large-scale experiments with over 1.8 trillion prompt tokens on 41 combinations of different corpora, retrievers, and backbone LLMs through the MEDRAG toolkit introduced in this work.Overall, MEDRAG improves the accuracy of six different LLMs by up to 18% over chain-of-thought prompting, elevating the performance of GPT-3.5 and Mixtral to GPT-4level.Our results show that the combination of various medical corpora and retrievers achieves the best performance.In addition, we discovered a log-linear scaling property and the "lostin-the-middle" effects in medical RAG.We believe our comprehensive evaluations can serve as practical guidelines for implementing RAG systems for medicine 1 .
Key Findings
1
Across six language models, MEDRAG improves accuracy by up to 18% over chain-of-thought prompting, raising GPT-3.5 and Mixtral to GPT-4-level performance.
2
Combining multiple medical corpora and retrievers achieves the strongest overall performance in medical RAG systems.
3
Medical RAG exhibits log-linear scaling and a “lost-in-the-middle” effect, revealing important retrieval-context limitations.
4
The MEDRAG toolkit enables large-scale evaluation of 41 combinations of medical corpora, retrievers, and backbone large language models.
5
The MIRAGE benchmark evaluates medical retrieval-augmented generation using 7,663 questions from five medical question-answering datasets.
Research Object
medical retrieval-augmented generation (RAG) systems for large language model question answering
Research Subject
the effects of corpora, retrievers, and backbone LLM configurations on medical QA accuracy, including scaling behavior and the lost-in-the-middle effect
Publication Details
Publication Date
2024-01-01
Journal
Publisher
ISSN
Cited by
257
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest