Red teaming ChatGPT in medicine to yield real-world insights on model behavior

Проверка ChatGPT методом red teaming в медицине для получения практических сведений о поведении модели
Yanan Ding, Archer Y. Yang, Jonathan H. Chen, Lü Tian, Matthew Schwede, Nigam H. Shah, Jesutofunmi A. Omiye, Haiwen Gui, Shawheen J. Rezaei, Roxana Daneshjou, Scott L. Fleming, Jason Fries, Balasubramanian Narasimhan, Juan M. Banda, Crystal Chang, Hodan Farah, Charbel Bou-Khalil, Ye-Jean Park, Akshay Swaminathan, Akaash Kolluri, Akash Chaurasia, Alejandro Lozano, Alice Heiman, Allison Sihan Jia, Amit Kaushal, Angela Y. Jia, Angelica Iacovelli, Arghavan Salles, Arpita Singhal, Benjamin Belai, Benjamin H. Jacobson, Binglan Li, Celeste H. Poe, Chandan Sanghera, Chenming Zheng, Conor Messer, Damien Varid Kettud, Deven Pandya, Dhamanpreet Kaur, Diana Hla, Diba Dindoust, Dominik Moehrle, Ross Duncan, Ellaine Chou, Eric Lin, Fateme Nateghi Haredasht, Cheng Ge, Irena Gao, Jacob Chang, Jake Silberg, Jiapeng Xu, J. Weston Jamison, John Tamaresis, Joshua Lazaro, Julie Lee, Karen Ebert Matthys, Kirsten R. Steffner, Luca Pegolotti, Malathi Srinivasan, Maniragav Manimaran, Minghe Zhang, Minh Hoai Nguyen, Mohsen Fathzadeh, Qian Zhao, Rika Bajra, Rohit Khurana, Ruhana Azam, R. W. Bartlett, Sang Truong, S. Varadha Raj, Solveig Behr, Sonia Onyeka, Sri Muppidi, Tarek Bandali, Tiffany Eulalio, Wenyuan Chen, Xuanyu Zhou, Ying Cui, Yuqi Tan, Yutong Liu
2025-03-07

clinical caseshallucinations and biaslarge language modelsmodel safetyred teaming
Red teaming, the practice of adversarially exposing unexpected or undesired model behaviors, is critical towards improving equity and accuracy of large language models, but non-model creator-affiliated red teaming is scant in healthcare. We convened teams of clinicians, medical and engineering students, and technical professionals (80 participants total) to stress-test models with real-world clinical cases and categorize inappropriate responses along axes of safety, privacy, hallucinations/accuracy, and bias. Six medically-trained reviewers re-analyzed prompt-response pairs and added qualitative annotations. Of 376 unique prompts (1504 responses), 20.1% were inappropriate (GPT-3.5: 25.8%; GPT-4.0: 16%; GPT-4.0 with Internet: 17.8%). Subsequently, we show the utility of our benchmark by testing GPT-4o, a model released after our event (20.4% inappropriate). 21.5% of responses appropriate with GPT-3.5 were inappropriate in updated models. We share insights for constructing red teaming prompts, and present our benchmark for iterative model assessments.
1
A non-model-creator-affiliated red-teaming effort engaged 80 clinicians, students, and technical professionals to evaluate medical LLM behavior using real-world clinical cases.
2
Among 376 unique prompts generating 1,504 responses, 20.1% were inappropriate across safety, privacy, hallucination/accuracy, and bias dimensions.
3
Inappropriate-response rates differed by model: 25.8% for GPT-3.5, 16% for GPT-4.0, and 17.8% for GPT-4.0 with Internet access.
4
Model updates did not consistently preserve safe behavior: 21.5% of responses appropriate with GPT-3.5 became inappropriate in updated models.
5
The benchmark detected 20.4% inappropriate responses in GPT-4o, despite its release after the original red-teaming event.

ChatGPT and related GPT models used for real-world clinical cases

Model behavior, including safety, privacy, hallucinations/accuracy, bias, and changes in inappropriate-response rates under adversarial clinical testing

Publication Details
Publication Date
2025-03-07
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Yanan Ding
Archer Y. Yang
Jonathan H. Chen
Lü Tian
Matthew Schwede
Nigam H. Shah
Jesutofunmi A. Omiye
Haiwen Gui
Shawheen J. Rezaei
Roxana Daneshjou
Scott L. Fleming
Jason Fries
Balasubramanian Narasimhan
Juan M. Banda
Crystal Chang
Hodan Farah
Charbel Bou-Khalil
Ye-Jean Park
Akshay Swaminathan
Akaash Kolluri
Akash Chaurasia
Alejandro Lozano
Alice Heiman
Allison Sihan Jia
Amit Kaushal
Angela Y. Jia
Angelica Iacovelli
Arghavan Salles
Arpita Singhal
Benjamin Belai
Benjamin H. Jacobson
Binglan Li
Celeste H. Poe
Chandan Sanghera
Chenming Zheng
Conor Messer
Damien Varid Kettud
Deven Pandya
Dhamanpreet Kaur
Diana Hla
Diba Dindoust
Dominik Moehrle
Ross Duncan
Ellaine Chou
Eric Lin
Fateme Nateghi Haredasht
Cheng Ge
Irena Gao
Jacob Chang
Jake Silberg
Jiapeng Xu
J. Weston Jamison
John Tamaresis
Joshua Lazaro
Julie Lee
Karen Ebert Matthys
Kirsten R. Steffner
Luca Pegolotti
Malathi Srinivasan
Maniragav Manimaran
Minghe Zhang
Minh Hoai Nguyen
Mohsen Fathzadeh
Qian Zhao
Rika Bajra
Rohit Khurana
Ruhana Azam
R. W. Bartlett
Sang Truong
S. Varadha Raj
Solveig Behr
Sonia Onyeka
Sri Muppidi
Tarek Bandali
Tiffany Eulalio
Wenyuan Chen
Xuanyu Zhou
Ying Cui
Yuqi Tan
Yutong Liu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%