Red teaming ChatGPT in medicine to yield real-world insights on model behavior
Проверка ChatGPT методом red teaming в медицине для получения практических сведений о поведении модели
2025-03-07
SCID: 54.1/dzzkv7ne
Discuss with AI
clinical caseshallucinations and biaslarge language modelsmodel safetyred teaming
Figures from the paper
Abstract (AI)
Red teaming, the practice of adversarially exposing unexpected or undesired model behaviors, is critical towards improving equity and accuracy of large language models, but non-model creator-affiliated red teaming is scant in healthcare. We convened teams of clinicians, medical and engineering students, and technical professionals (80 participants total) to stress-test models with real-world clinical cases and categorize inappropriate responses along axes of safety, privacy, hallucinations/accuracy, and bias. Six medically-trained reviewers re-analyzed prompt-response pairs and added qualitative annotations. Of 376 unique prompts (1504 responses), 20.1% were inappropriate (GPT-3.5: 25.8%; GPT-4.0: 16%; GPT-4.0 with Internet: 17.8%). Subsequently, we show the utility of our benchmark by testing GPT-4o, a model released after our event (20.4% inappropriate). 21.5% of responses appropriate with GPT-3.5 were inappropriate in updated models. We share insights for constructing red teaming prompts, and present our benchmark for iterative model assessments.
Key Findings
1
A non-model-creator-affiliated red-teaming effort engaged 80 clinicians, students, and technical professionals to evaluate medical LLM behavior using real-world clinical cases.
2
Among 376 unique prompts generating 1,504 responses, 20.1% were inappropriate across safety, privacy, hallucination/accuracy, and bias dimensions.
3
Inappropriate-response rates differed by model: 25.8% for GPT-3.5, 16% for GPT-4.0, and 17.8% for GPT-4.0 with Internet access.
4
Model updates did not consistently preserve safe behavior: 21.5% of responses appropriate with GPT-3.5 became inappropriate in updated models.
5
The benchmark detected 20.4% inappropriate responses in GPT-4o, despite its release after the original red-teaming event.
Research Object
ChatGPT and related GPT models used for real-world clinical cases
Research Subject
Model behavior, including safety, privacy, hallucinations/accuracy, bias, and changes in inappropriate-response rates under adversarial clinical testing
Publication Details
Publication Date
2025-03-07
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest