AgentClinic: a multimodal benchmark for tool-using clinical AI agents

AgentClinic: мультимодальный бенчмарк для клинических ИИ-агентов, использующих инструменты
Jae‐Joong Kim, Eduardo Pontes Reis, Samuel Schmidgall, Rojin Ziaei, Carl Harris, Jeffrey Jopling, Michael Moor
2026-04-27

clinical simulation benchmarkelectronic health recordsmultimodal clinical AI agentssequential decision-makingtool use
Evaluating large language models (LLM) in clinical scenarios is crucial to assessing their potential clinical utility. Existing benchmarks rely heavily on static question-answering, which does not accurately depict the complex, sequential nature of clinical decision-making. Here, we introduce AgentClinic, a multimodal agent benchmark for evaluating LLMs in simulated clinical environments that include patient interactions, multimodal data collection under incomplete information, and the usage of various tools, resulting in an in-depth evaluation across nine medical specialties and seven languages. We find that solving MedQA problems in the sequential decision-making format of AgentClinic is considerably more challenging, resulting in diagnostic accuracies that can drop to below a tenth of the original accuracy. Overall, we observe that agents sourced from Claude-3.5 outperform other LLM backbones in most settings. Nevertheless, we see stark differences in the LLMs' ability to make use of tools, such as experiential learning, adaptive retrieval, and reflection cycles. Strikingly, Llama-3 shows up to 92% relative improvements with the notebook tool that allows for writing and editing notes that persist across cases. To further scrutinize our clinical simulations, we leverage real-world electronic health records, perform a clinical reader study, perturb agents with biases, and explore patient-centric metrics that this interactive environment enables.
1
AgentClinic introduces a multimodal benchmark for tool-using clinical agents in simulated environments with patient interaction, incomplete information, and sequential tool use.
2
Claude-3.5-based agents outperform other evaluated LLM backbones in most settings, while models differ markedly in their ability to use experiential learning, adaptive retrieval, and reflection tools.
3
Llama-3 achieves up to 92% relative improvement when using a persistent notebook tool for writing and editing notes across cases.
4
Recasting MedQA as sequential decision-making substantially increases difficulty, with diagnostic accuracy dropping below one-tenth of the original accuracy.
5
The benchmark evaluates clinical agents across nine medical specialties and seven languages, extending beyond static question-answering tasks.
6
The benchmark is further examined using real-world electronic health records, clinical reader studies, bias perturbations, and patient-centric metrics enabled by interactive simulations.

LLM-based clinical AI agents operating in simulated multimodal clinical environments

their sequential clinical decision-making, diagnostic accuracy, tool use, and patient-centric performance under incomplete information

Publication Details
Publication Date
2026-04-27
Journal
Publisher
ISSN
Cited by
7
Access Type
Author Information
Authors
Jae‐Joong Kim
Eduardo Pontes Reis
Samuel Schmidgall
Rojin Ziaei
Carl Harris
Jeffrey Jopling
Michael Moor
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%