An autonomous multimodal AI agent for evidence-grounded ophthalmic diagnosis

Автономный мультимодальный ИИ-агент для офтальмологической диагностики на основе доказательств
Kaikai Zhao, Qixuan Sun, Daohuan Kang, Tao Yu, Wenzheng Han, Rui Yao, Rupesh Agrawal, Gui-shuang Ying, Andrzej Grzybowski, Kai Jin
2026-08-01

B-scan ultrasonographyautonomous AI agentevidence-grounded reportingfundus photographymultimodal ophthalmic diagnosis
Multimodal ophthalmic diagnosis requires integrating fundus photography, B-scan ultrasonography, and medical evidence, yet most artificial intelligence (AI) systems remain single-task or weakly grounded. AgentEYE is an auditable multimodal agent that routes ocular images to specialized fundus and B-scan tools, retrieves guideline/web evidence, and synthesizes evidence-grounded reports. In a 302-case internal benchmark, AgentEYE shows higher diagnostic correctness and completeness than large language model (LLM)-only baselines and an ablation without specialized imaging tools; performance remains similar to the no-retrieval ablation, indicating that retrieval mainly supports evidence grounding and citation auditability. Blinded evaluation of 200 cases by three ophthalmologists confirms improved diagnostic correctness, completeness, safety, and citation grounding versus an LLM-only self-citation baseline. External analyses show distribution-dependent performance. These findings support AgentEYE as a traceable decision-support prototype requiring prospective multicenter validation.
1
AgentEYE integrates fundus photography, B-scan ultrasonography, and retrieved medical evidence through specialized tools to generate auditable ophthalmic diagnostic reports.
2
Blinded assessment of 200 cases by three ophthalmologists found AgentEYE improved diagnostic correctness, completeness, safety, and citation grounding compared with an LLM-only self-citation baseline.
3
External analyses revealed distribution-dependent performance, supporting the need for prospective multicenter validation before clinical use.
4
On a 302-case internal benchmark, AgentEYE achieved higher diagnostic correctness and completeness than LLM-only baselines and a version without specialized imaging tools.
5
Removing evidence retrieval produced similar diagnostic performance, indicating retrieval primarily improves evidence grounding and citation auditability rather than diagnosis.

AgentEYE, an autonomous multimodal AI agent for ophthalmic diagnosis using fundus photography, B-scan ultrasonography, and medical evidence

evidence-grounded ophthalmic diagnostic performance, including diagnostic correctness, completeness, safety, traceability, and citation grounding

Publication Details
Publication Date
2026-08-01
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Kaikai Zhao
Qixuan Sun
Daohuan Kang
Tao Yu
Wenzheng Han
Rui Yao
Rupesh Agrawal
Gui-shuang Ying
Andrzej Grzybowski
Kai Jin
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%