Retrieval augmented generation for large language models in healthcare: A systematic review
Генерация с дополнением поиском для больших языковых моделей в здравоохранении: систематический обзор
2025-06-11
SCID: 54.1/xep4rurg
Discuss with AI
ethical considerationsevaluation frameworkshealthcarelarge language modelsretrieval-augmented generation
Figures from the paper
Abstract (AI)
Large Language Models (LLMs) have demonstrated promising capabilities to solve complex tasks in critical sectors such as healthcare. However, LLMs are limited by their training data which is often outdated, the tendency to generate inaccurate ("hallucinated") content and a lack of transparency in the content they generate. To address these limitations, retrieval augmented generation (RAG) grounds the responses of LLMs by exposing them to external knowledge sources. However, in the healthcare domain there is currently a lack of systematic understanding of which datasets, RAG methodologies and evaluation frameworks are available. This review aims to bridge this gap by assessing RAG-based approaches employed by LLMs in healthcare, focusing on the different steps of retrieval, augmentation and generation. Additionally, we identify the limitations, strengths and gaps in the existing literature. Our synthesis shows that 78.9% of studies used English datasets and 21.1% of the datasets are in Chinese. We find that a range of techniques are employed RAG-based LLMs in healthcare, including Naive RAG, Advanced RAG, and Modular RAG. Surprisingly, proprietary models such as GPT-3.5/4 are the most used for RAG applications in healthcare. We find that there is a lack of standardised evaluation frameworks for RAG-based applications. In addition, the majority of the studies do not assess or address ethical considerations related to RAG in healthcare. It is important to account for ethical challenges that are inherent when AI systems are implemented in the clinical setting. Lastly, we highlight the need for further research and development to ensure responsible and effective adoption of RAG in the medical domain.
Key Findings
1
Among reviewed studies, 78.9% used English datasets and 21.1% used Chinese datasets, indicating limited linguistic diversity.
2
Healthcare RAG systems employ Naive RAG, Advanced RAG, and Modular RAG approaches, with proprietary GPT-3.5/4 models used most frequently.
3
Most studies do not evaluate ethical considerations, highlighting the need for responsible clinical deployment and further research.
4
The literature lacks standardized evaluation frameworks for healthcare RAG applications, limiting consistent comparison of system performance.
5
The systematic review identifies datasets, retrieval methods, augmentation strategies, and evaluation frameworks used in healthcare RAG applications.
Research Object
Retrieval augmented generation (RAG)-based approaches applied to large language models in healthcare
Research Subject
the available datasets, retrieval–augmentation–generation methodologies, evaluation frameworks, strengths, limitations, gaps, and ethical considerations of RAG-based healthcare applications
Publication Details
Publication Date
2025-06-11
Journal
Publisher
ISSN
Cited by
219
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai8
Large language models encode clinical knowledge2023
Sparks of Artificial General Intelligence: Early experiments with GPT-42023
Benchmarking Retrieval-Augmented Generation for Medicine2024
The Power of Noise: Redefining Retrieval for RAG Systems2024
Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model2024
Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer2019
XGBoost2016
The PRISMA Statement for Reporting Systematic Reviews and Meta-Analyses of Studies That Evaluate Health Care Interventions: Explanation and Elaboration2009