Evaluation of large language models as a diagnostic aid for complex medical cases
Оценка больших языковых моделей как вспомогательного инструмента для диагностики сложных медицинских случаев
2024-06-20
SCID: 54.1/9wzyb6qm
Discuss with AI
GPT-4Jaccard Similarity Indexcomplex clinical casesdifferential diagnosislarge language models
Figures from the paper
Abstract (AI)
Background The use of large language models (LLM) has recently gained popularity in diverse areas, including answering questions posted by patients as well as medical professionals. Objective To evaluate the performance and limitations of LLMs in providing the correct diagnosis for a complex clinical case. Design Seventy-five consecutive clinical cases were selected from the Massachusetts General Hospital Case Records, and differential diagnoses were generated by OpenAI’s GPT3.5 and 4 models. Results The mean number of diagnoses provided by the Massachusetts General Hospital case discussants was 16.77, by GPT3.5 30 and by GPT4 15.45 ( p < 0.0001). GPT4 was more frequently able to list the correct diagnosis as first (22% versus 20% with GPT3.5, p = 0.86), provide the correct diagnosis among the top three generated diagnoses (42% versus 24%, p = 0.075). GPT4 was better at providing the correct diagnosis, when the different diagnoses were classified into groups according to the medical specialty and include the correct diagnosis at any point in the differential list (68% versus 48%, p = 0.0063). GPT4 provided a differential list that was more similar to the list provided by the case discussants than GPT3.5 (Jaccard Similarity Index 0.22 versus 0.12, p = 0.001). Inclusion of the correct diagnosis in the generated differential was correlated with PubMed articles matching the diagnosis (OR 1.40, 95% CI 1.25–1.56 for GPT3.5, OR 1.25, 95% CI 1.13–1.40 for GPT4), but not with disease incidence. Conclusions and relevance The GPT4 model was able to generate a differential diagnosis list with the correct diagnosis in approximately two thirds of cases, but the most likely diagnosis was often incorrect for both models. In its current state, this tool can at most be used as an aid to expand on potential diagnostic considerations for a case, and future LLMs should be trained which account for the discrepancy between disease incidence and availability in the literature.
Key Findings
1
Correct-diagnosis inclusion correlated with the availability of matching PubMed literature, but not with disease incidence, suggesting literature availability influences model outputs.
2
GPT-4 generated shorter differentials than GPT-3.5 and more closely matched Massachusetts General Hospital discussants’ lists, with Jaccard similarity indices of 0.22 versus 0.12.
3
GPT-4 included the correct diagnosis somewhere in its differential list in 68% of cases, compared with 48% for GPT-3.5 when diagnoses were grouped by medical specialty.
4
The correct diagnosis was ranked first in only 22% of cases by GPT-4 and 20% by GPT-3.5, indicating frequent failure to identify the most likely diagnosis.
5
The models may currently serve only as aids for broadening diagnostic considerations, rather than as reliable systems for selecting the primary diagnosis.
Research Object
GPT3.5 and GPT4 large language models applied to 75 complex clinical cases from the Massachusetts General Hospital Case Records
Research Subject
Diagnostic performance and limitations, including the accuracy, ranking, composition, and similarity of generated differential diagnoses
Publication Details
Publication Date
2024-06-20
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest