OneEHR: Reproducible and AI Agent-Ready Longitudinal EHR Analysis Toolkit
OneEHR: воспроизводимый инструментарий для анализа продольных данных ЭМК, готовый к использованию ИИ-агентами
2026-08-07
SCID: 54.1/4a3a5pcr
Discuss with AI
AI agent-ready modelingOneEHR toolkitclinical AI experimentslongitudinal EHR analysisreproducible EHR research
Figures from the paper
Abstract (AI)
Electronic health records support a wide spectrum of clinical prediction and decision-support studies, but reproducible EHR research now requires more than training a single predictive model. As the field expands from machine learning and deep learning to LLM-based and agentic AI, differences in cohort construction, temporal preprocessing, label definitions, patient-level splits, and evaluation protocols can overshadow the methods being compared, making fair comparison and model selection difficult in practice. This tutorial presents OneEHR, an open-source toolkit that defines a unified experiment contract for modern EHR modeling and enables head-to-head comparison among conventional, neural, LLM-based, and agentic methods through a single configuration-driven interface. The three-hour hands-on session interleaves a methodological survey with guided practice: participants will learn why EHR experiments are vulnerable to leakage, distribution shift, and irreproducible preprocessing, and then use OneEHR to configure, execute, compare, and interpret experiments across this method spectrum. Attendees will leave with reusable configurations and a practical framework for integrating reproducible workflows into their own clinical AI research. Code and documentation are available at https://medx-pku.github.io/OneEHR/.
Key Findings
1
OneEHR is an open-source toolkit providing a unified experiment contract for reproducible longitudinal EHR modeling.
2
OneEHR provides reusable configurations and workflows for executing, comparing, and interpreting clinical AI experiments.
3
The configuration-driven interface enables head-to-head comparison of conventional, neural, LLM-based, and agentic AI methods.
4
The toolkit standardizes cohort construction, temporal preprocessing, label definitions, patient-level splitting, and evaluation protocols.
5
The tutorial highlights leakage, distribution shift, and irreproducible preprocessing as major threats to fair EHR method comparison.
Research Object
longitudinal electronic health record (EHR) analysis and modeling experiments
Research Subject
reproducibility and fair comparative evaluation of conventional, neural, LLM-based, and agentic AI methods under standardized cohort construction, temporal preprocessing, labeling, splitting, and evaluation protocols
Publication Details
Publication Date
2026-08-07
Journal
Publisher
ISSN
Cited by
0
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest