MAIA: A Multidimensional Benchmark for Assessing Medical AI Agents

MAIA: многомерный бенчмарк для оценки медицинских ИИ-агентов
Xin Ding, Jiaxin Ding, Yule Xie, Xinbing Wang, Biao Peng, Yan Li, Yichen Li, Andrew Huang, Yinan Wang
2026-04-21

biomedical knowledge graphsclinical-pathway reasoningmedical AI agentsmulti-hop reasoningmultidimensional benchmark
Large language models show remarkable potential in medical scenarios, especially as autonomous agents for complex clinical reasoning. Rigorous evaluation is essential to ensure their reliability in real-world healthcare applications. However, existing medical benchmarks suffer from narrow task scopes, dependence on public datasets prone to data leakage, and limited coverage of diverse agent capabilities. To address these gaps, we introduce Medical AI Assessment (MAIA), a comprehensive benchmark evaluating medical agents along three dimensions: retrieval-based medical questions generated through biomedical APIs, multi-hop reasoning tasks derived from curated biomedical knowledge graphs, clinical-pathway reasoning questions constructed from authoritative guidelines. MAIA leverages large language models for automatic question generation, reducing manual effort while maintaining clinical fidelity and reasoning depth. Experiments across base and reasoning models reveal both strengths and gaps, underscoring MAIA’s value for advancing medical agent evaluation. MAIA is publicly available at https://huggingface.co/datasets/DiligentDing/MAIA.
1
Experiments with base and reasoning models identify both capabilities and shortcomings, demonstrating the benchmark’s usefulness for comprehensive medical-agent evaluation.
2
Large language models automate question generation while preserving clinical fidelity and reasoning depth, reducing the manual effort required for benchmark construction.
3
MAIA is a multidimensional benchmark for evaluating medical AI agents across retrieval, multi-hop biomedical reasoning, and clinical-pathway reasoning.
4
MAIA is publicly released, enabling accessible assessment and comparison of medical AI agents.
5
The benchmark generates retrieval questions through biomedical APIs, multi-hop tasks from curated biomedical knowledge graphs, and clinical questions from authoritative guidelines.

medical AI agents

their reliability and capabilities in retrieval-based medical question answering, multi-hop biomedical reasoning, and clinical-pathway reasoning

Publication Details
Publication Date
2026-04-21
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Xin Ding
Jiaxin Ding
Yule Xie
Xinbing Wang
Biao Peng
Yan Li
Yichen Li
Andrew Huang
Yinan Wang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%