MedAgentBench: A Virtual EHR Environment to Benchmark Medical LLM Agents

MedAgentBench: виртуальная среда электронной медицинской карты для тестирования медицинских агентов на основе больших языковых моделей
Yixing Jiang, Kameron Collin Black, Gloria Geng, Dae-Gyun Park, James Zou, Andrew Y. Ng, Jonathan H. Chen
2025-08-14

FHIR-compliant environmentMedAgentBenchagent-oriented benchmarkelectronic health recordsmedical LLM agents
BACKGROUND Recent large language models (LLMs) have demonstrated significant advancements, particularly in their ability to serve as agents, thereby surpassing their traditional role as chatbots.These agents can leverage their planning and tool utilization capabilities to address tasks specified at a high level.This suggests new potential to reduce the burden of administrative tasks and address current health care staff shortages.However, a standardized dataset to benchmark the agent capabilities of LLMs in medical applications is currently lacking, making it difficult to evaluate their performance on complex tasks in interactive health care environments. METHODSTo address this gap in the deployment of agentic artificial intelligence (AI) in health care, we introduce MedAgentBench, a broad evaluation suite designed to assess the agent capabilities of LLMs within medical records contexts.MedAgentBench encompasses 300 patient-specific clinically derived tasks from 10 categories written by human physicians, realistic profiles of 100 patients with over 700,000 data elements, a Fast Healthcare Interoperability Resources-compliant interactive environment, and an accompanying codebase.The environment uses standard application programming interfaces and communication infrastructure used in modern electronic health record (EHR) systems so that it can be easily migrated into live EHR systems. RESULTSMedAgentBench presents an unsaturated agent-oriented benchmark at which current state-of-the-art LLMs exhibit some ability to succeed.The best model (Claude 3.5 Sonnet v2) achieves a success rate of 69.67%.However, there is still substantial room for improvement, which gives the community a clear direction for future optimization efforts.Furthermore, there is significant variation in performance across task categories.CONCLUSIONS Agent-based task frameworks and benchmarks are the necessary next step to advance the potential and capabilities for effectively improving and integrating AI systems into clinical workflows.MedAgentBench establishes this and is publicly available at https://github .com /stanfordmlgroup /MedAgentBench, offering a valuable framework for model developers to track progress and drive continuous improvements in the agent capabilities of LLMs within the medical domain.
1
LLM-agent performance varies significantly across task categories, highlighting uneven capabilities in medical EHR interactions.
2
MedAgentBench introduces a standardized benchmark for evaluating medical LLM agents on complex, interactive electronic health record tasks.
3
MedAgentBench provides a FHIR-compliant interactive environment using standard EHR APIs and communication infrastructure, supporting potential migration into live systems.
4
The benchmark contains 300 physician-written, patient-specific tasks across 10 categories, realistic profiles for 100 patients, and more than 700,000 data elements.
5
The benchmark remains unsaturated: Claude 3.5 Sonnet v2 achieves the highest reported success rate at 69.67%, indicating substantial room for improvement.

MedAgentBench’s virtual, FHIR-compliant interactive electronic health record environment containing realistic patient profiles and clinically derived tasks

The task-solving capabilities and performance of medical LLM agents in interactive EHR contexts, including variation across task categories

Publication Details
Publication Date
2025-08-14
Journal
Publisher
ISSN
Cited by
55
Access Type
Author Information
Authors
Yixing Jiang
Kameron Collin Black
Gloria Geng
Dae-Gyun Park
James Zou
Andrew Y. Ng
Jonathan H. Chen
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%