Clinical entity augmented retrieval for clinical information extraction

Извлечение с расширением клиническими сущностями для извлечения клинической информации
Jonathan H. Chen, Nigam H. Shah, Akshay Swaminathan, Fateme Nateghi Haredasht, Iván López, P. Stephen, Karthik S. Vedula, Sanjana Narayanan, April S. Liang, Steven Tate, Manoj Maddali, Robert J. Gallo
2025-01-19

Clinical Entity Augmented Retrievalclinical information extractionembedding-based retrievallarge language modelsretrieval-augmented generation
Large language models (LLMs) with retrieval-augmented generation (RAG) have improved information extraction over previous methods, yet their reliance on embeddings often leads to inefficient retrieval. We introduce CLinical Entity Augmented Retrieval (CLEAR), a RAG pipeline that retrieves information using entities. We compared CLEAR to embedding RAG and full-note approaches for extracting 18 variables using six LLMs across 20,000 clinical notes. Average F1 scores were 0.90, 0.86, and 0.79; inference times were 4.95, 17.41, and 20.08 s per note; average model queries were 1.68, 4.94, and 4.18 per note; and average input tokens were 1.1k, 3.8k, and 6.1k per note for CLEAR, embedding RAG, and full-note approaches, respectively. In conclusion, CLEAR utilizes clinical entities for information retrieval and achieves >70% reduction in token usage and inference time with improved performance compared to modern methods.
1
Across 20,000 clinical notes, six LLMs, and 18 variables, CLEAR achieved the highest average F1 score: 0.90 versus 0.86 for embedding RAG and 0.79 for full-note processing.
2
CLEAR is a clinical entity-based retrieval-augmented generation pipeline for clinical information extraction.
3
CLEAR reduced average inference time to 4.95 seconds per note, compared with 17.41 seconds for embedding RAG and 20.08 seconds for full-note approaches.
4
CLEAR required fewer model queries per note, averaging 1.68 versus 4.94 for embedding RAG and 4.18 for full-note processing.
5
CLEAR used 1.1k average input tokens per note, over 70% fewer than embedding RAG (3.8k) and full-note approaches (6.1k), while improving extraction performance.

clinical notes processed for clinical information extraction

the effectiveness and efficiency of entity-based retrieval-augmented generation for extracting 18 clinical variables, including accuracy, inference time, model queries, and token usage

Publication Details
Publication Date
2025-01-19
Journal
Publisher
ISSN
Cited by
67
Access Type
Author Information
Authors
Jonathan H. Chen
Nigam H. Shah
Akshay Swaminathan
Fateme Nateghi Haredasht
Iván López
P. Stephen
Karthik S. Vedula
Sanjana Narayanan
April S. Liang
Steven Tate
Manoj Maddali
Robert J. Gallo
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%