Bias in medical AI: Implications for clinical decision-making
Предвзятость в медицинском искусственном интеллекте: последствия для принятия клинических решений
2024-11-07
SCID: 54.1/ffh7jb9d
Discuss with AI
algorithmic biasclinical decision-makinghealthcare disparitiesmedical artificial intelligencemodel interpretability
Figures from the paper
Abstract (AI)
Biases in medical artificial intelligence (AI) arise and compound throughout the AI lifecycle. These biases can have significant clinical consequences, especially in applications that involve clinical decision-making. Left unaddressed, biased medical AI can lead to substandard clinical decisions and the perpetuation and exacerbation of longstanding healthcare disparities. We discuss potential biases that can arise at different stages in the AI development pipeline and how they can affect AI algorithms and clinical decision-making. Bias can occur in data features and labels, model development and evaluation, deployment, and publication. Insufficient sample sizes for certain patient groups can result in suboptimal performance, algorithm underestimation, and clinically unmeaningful predictions. Missing patient findings can also produce biased model behavior, including capturable but nonrandomly missing data, such as diagnosis codes, and data that is not usually or not easily captured, such as social determinants of health. Expertly annotated labels used to train supervised learning models may reflect implicit cognitive biases or substandard care practices. Overreliance on performance metrics during model development may obscure bias and diminish a model's clinical utility. When applied to data outside the training cohort, model performance can deteriorate from previous validation and can do so differentially across subgroups. How end users interact with deployed solutions can introduce bias. Finally, where models are developed and published, and by whom, impacts the trajectories and priorities of future medical AI development. Solutions to mitigate bias must be implemented with care, which include the collection of large and diverse data sets, statistical debiasing methods, thorough model evaluation, emphasis on model interpretability, and standardized bias reporting and transparency requirements. Prior to real-world implementation in clinical settings, rigorous validation through clinical trials is critical to demonstrate unbiased application. Addressing biases across model development stages is crucial for ensuring all patients benefit equitably from the future of medical AI.
Key Findings
1
Bias can arise and compound across the entire medical AI lifecycle, including data, labeling, modeling, evaluation, deployment, and publication.
2
Expert annotations may encode implicit cognitive biases or substandard care practices, while reliance on aggregate performance metrics can obscure clinically important bias.
3
Missing or nonrandomly captured findings, including diagnosis codes and social determinants of health, can create systematically biased model behavior.
4
Model performance may deteriorate outside the training cohort and degrade differentially across patient subgroups, with user interaction introducing additional bias.
5
The paper recommends large diverse datasets, statistical debiasing, thorough subgroup evaluation, interpretability, and standardized bias reporting before clinical implementation.
6
Underrepresentation of patient groups can reduce model performance, cause algorithmic underestimation, and produce clinically meaningless predictions.
Research Object
medical artificial intelligence systems used in clinical decision-making
Research Subject
biases throughout the AI lifecycle and their effects on algorithm performance, clinical utility, decision-making, and healthcare disparities
Publication Details
Publication Date
2024-11-07
Journal
Publisher
ISSN
Cited by
543
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest