A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation
Платформа для оценки клинической безопасности и частоты галлюцинаций больших языковых моделей при суммировании медицинских текстов
2025-05-13
SCID: 54.1/yd3ty9jn
Discuss with AI
clinical note generationclinical safetyerror taxonomyhallucination ratesmedical text summarisation
Figures from the paper
Abstract (AI)
Integrating large language models (LLMs) into healthcare can enhance workflow efficiency and patient care by automating tasks such as summarising consultations. However, the fidelity between LLM outputs and ground truth information is vital to prevent miscommunication that could lead to compromise in patient safety. We propose a framework comprising (1) an error taxonomy for classifying LLM outputs, (2) an experimental structure for iterative comparisons in our LLM document generation pipeline, (3) a clinical safety framework to evaluate the harms of errors, and (4) a graphical user interface, CREOLA, to facilitate these processes. Our clinical error metrics were derived from 18 experimental configurations involving LLMs for clinical note generation, consisting of 12,999 clinician-annotated sentences. We observed a 1.47% hallucination rate and a 3.45% omission rate. By refining prompts and workflows, we successfully reduced major errors below previously reported human note-taking rates, highlighting the framework's potential for safer clinical documentation.
Key Findings
1
Clinical error metrics were derived from 18 LLM clinical-note-generation configurations covering 12,999 clinician-annotated sentences.
2
Prompt and workflow refinement reduced major errors below previously reported rates for human clinical note-taking.
3
The evaluated systems exhibited a 1.47% hallucination rate and a 3.45% omission rate.
4
The framework supports systematic identification and mitigation of LLM errors to enable safer automated clinical documentation.
5
The paper introduces a framework for evaluating LLM clinical summarization safety, combining an error taxonomy, iterative experiments, clinical harm assessment, and the CREOLA interface.
Research Object
LLM-generated medical text summarisation and clinical note generation
Research Subject
Clinical safety, hallucination, omission, and error rates in LLM-generated clinical documentation
Publication Details
Publication Date
2025-05-13
Journal
Publisher
ISSN
Cited by
341
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest