A large language model-driven multidisciplinary AI agent system predicts delirium in emergency critically ill patients

Мультидисциплинарная система ИИ-агентов на основе большой языковой модели прогнозирует делирий у экстренных пациентов в критическом состоянии
Wen Shang, Tongyue Shi, Qingbian Ma, Guilan Kong
2026-08-01

MIMIC-IVdelirium risk predictionlarge language modelmulti-agent systemretrieval-augmented generation
Delirium occurs frequently in emergency departments and is associated with poor outcomes and increased burden. Early delirium risk prediction is crucial for timely prevention and intervention in emergency care, but most existing models focus on intensive care unit (ICU) populations and offer limited interpretability and interactivity. We propose DeLiriuMAgents, a large language model (LLM)-driven multi-agent system for predicting delirium risk in emergency critically ill patients. It simulates multidisciplinary clinical consultation by integrating data-driven, machine learning-based risk prediction; LLM-based virtual specialist reasoning in emergency medicine, neurology, and psychiatry; and medical evidence via retrieval-augmented generation to reach a final decision. In model development, Medical Information Mart for Intensive Care (MIMIC)-IV is used for model derivation and internal validation; a multicenter Peking University (PKU) cohort from two hospitals in China and the eICU Collaborative Research Database (eICU-CRD) cohort are used for external validation. It achieves accuracy/sensitivity/specificity of 0.749/0.762/0.747, 0.731/0.708/0.736, and 0.670/0.708/0.665 on MIMIC-IV, PKU, and eICU-CRD validation sets, respectively. Chart review and clinician evaluation verify the interpretability and usefulness of its reports.
1
Chart review and clinician evaluation supported the interpretability and practical usefulness of the system’s generated reports.
2
DeLiriuMAgents is an LLM-driven multidisciplinary multi-agent system designed to predict delirium risk in emergency critically ill patients.
3
It was developed and internally validated using MIMIC-IV, with external validation in a multicenter Peking University cohort and the eICU-CRD.
4
The system combines machine-learning risk prediction, virtual specialist reasoning in emergency medicine, neurology and psychiatry, and retrieval-augmented medical evidence.
5
Validation performance was accuracy/sensitivity/specificity of 0.749/0.762/0.747 on MIMIC-IV, 0.731/0.708/0.736 on PKU, and 0.670/0.708/0.665 on eICU-CRD.

delirium risk in emergency critically ill patients

prediction accuracy, interpretability, and clinical usefulness of delirium risk assessment

Publication Details
Publication Date
2026-08-01
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Wen Shang
Tongyue Shi
Qingbian Ma
Guilan Kong
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%