Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines

Улучшение применения больших языковых моделей в биомедицине с помощью генерации с дополнением извлечённой информацией: систематический обзор, метаанализ и рекомендации по клинической разработке
Allison B. McCoy, Adam Wright, Siru Liu
2025-01-03

biomedicineclinical development guidelineslarge language modelsretrieval-augmented generationsystematic review and meta-analysis
OBJECTIVE: The objectives of this study are to synthesize findings from recent research of retrieval-augmented generation (RAG) and large language models (LLMs) in biomedicine and provide clinical development guidelines to improve effectiveness. MATERIALS AND METHODS: We conducted a systematic literature review and a meta-analysis. The report was created in adherence to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 analysis. Searches were performed in 3 databases (PubMed, Embase, PsycINFO) using terms related to "retrieval augmented generation" and "large language model," for articles published in 2023 and 2024. We selected studies that compared baseline LLM performance with RAG performance. We developed a random-effect meta-analysis model, using odds ratio as the effect size. RESULTS: Among 335 studies, 20 were included in this literature review. The pooled effect size was 1.35, with a 95% confidence interval of 1.19-1.53, indicating a statistically significant effect (P = .001). We reported clinical tasks, baseline LLMs, retrieval sources and strategies, as well as evaluation methods. DISCUSSION: Building on our literature review, we developed Guidelines for Unified Implementation and Development of Enhanced LLM Applications with RAG in Clinical Settings to inform clinical applications using RAG. CONCLUSION: Overall, RAG implementation showed a 1.35 odds ratio increase in performance compared to baseline LLMs. Future research should focus on (1) system-level enhancement: the combination of RAG and agent, (2) knowledge-level enhancement: deep integration of knowledge into LLM, and (3) integration-level enhancement: integrating RAG systems within electronic health records.
1
A systematic review screened 335 studies and included 20 comparing baseline large language models with retrieval-augmented generation in biomedicine.
2
Future priorities include combining RAG with agents, deeply integrating knowledge into LLMs, and embedding RAG systems within electronic health records.
3
Meta-analysis found that RAG improved performance over baseline LLMs, with a pooled odds ratio of 1.35 (95% CI, 1.19–1.53; P = .001).
4
The authors developed guidelines for unified implementation and development of enhanced RAG-based LLM applications in clinical settings.
5
The review characterized clinical tasks, baseline models, retrieval sources and strategies, and evaluation methods used in biomedical RAG studies.

Retrieval-augmented generation (RAG) applications using large language models (LLMs) in biomedicine and clinical settings

The effectiveness and performance improvement of RAG-enhanced LLMs compared with baseline LLMs across biomedical and clinical tasks

Publication Details
Publication Date
2025-01-03
Journal
Publisher
ISSN
Cited by
166
Access Type
Author Information
Authors
Allison B. McCoy
Adam Wright
Siru Liu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%