Evaluating large language models as agents in the clinic

Оценка больших языковых моделей как агентов в клинической практике
Atul J. Butte, Madhumita Sushil, Brenda Y. Miao, Nikita Mehandru, Eduardo Rodriguez Almaraz, Ahmed M. Alaa
2024-04-03

Artificial Intelligence Structured Clinical ExaminationsLLM agentsclinical decision supportclinical workflow evaluationlarge language models
Recent developments in large language models (LLMs) have unlocked opportunities for healthcare, from information synthesis to clinical decision support. These LLMs are not just capable of modeling language, but can also act as intelligent “agents” that interact with stakeholders in open-ended conversations and even influence clinical decision-making. Rather than relying on benchmarks that measure a model’s ability to process clinical data or answer standardized test questions, LLM agents can be modeled in high-fidelity simulations of clinical settings and should be assessed for their impact on clinical workflows. These evaluation frameworks, which we refer to as “Artificial Intelligence Structured Clinical Examinations” (“AI-SCE”), can draw from comparable technologies where machines operate with varying degrees of self-governance, such as self-driving cars, in dynamic environments with multiple stakeholders. Developing these robust, real-world clinical evaluations will be crucial towards deploying LLM agents in medical settings.
1
AI-SCE frameworks should evaluate agents in dynamic, multi-stakeholder environments, drawing methodological inspiration from self-driving-car assessment.
2
Conventional benchmarks focused on clinical data processing or standardized questions may not adequately evaluate LLM agents’ effects on real clinical workflows.
3
LLMs can function as intelligent clinical agents that engage stakeholders through open-ended conversations and potentially influence clinical decision-making.
4
Robust, real-world clinical evaluations are identified as essential before deploying LLM agents in medical settings.
5
The paper proposes Artificial Intelligence Structured Clinical Examinations (AI-SCE) using high-fidelity simulations of clinical settings to assess agent performance and impact.

large language model agents in clinical settings

their impact on clinical workflows and the development of high-fidelity, real-world evaluation frameworks for deployment in medical settings

Publication Details
Publication Date
2024-04-03
Journal
Publisher
ISSN
Cited by
139
Access Type
Author Information
Authors
Atul J. Butte
Madhumita Sushil
Brenda Y. Miao
Nikita Mehandru
Eduardo Rodriguez Almaraz
Ahmed M. Alaa
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%