Risks of AI scientists: prioritizing safeguarding over autonomy

Риски «AI-учёных»: приоритет защиты над автономией
Arman Cohan, Zhiyong Lu, Qiao Jin, Mark Gerstein, Zhuosheng Zhang, Meng Qu, Xiangru Tang, Wangchunshu Zhou, Tongxin Yuan, Kunlun Zhu, Yichi Zhang, Yilun Zhao, J. Tang, Dov Greenbaum
2025-09-18

AI scientistsagent alignmenthuman regulationlarge language modelssafeguarding vs autonomy
AI scientists powered by large language models have demonstrated substantial promise in autonomously conducting experiments and facilitating scientific discoveries across various disciplines. While their capabilities are promising, these agents also introduce novel vulnerabilities that require careful consideration for safety. However, there has been limited comprehensive exploration of these vulnerabilities. This perspective examines vulnerabilities in AI scientists, shedding light on potential risks associated with their misuse, and emphasizing the need for safety measures. We begin by providing an overview of the potential risks inherent to AI scientists, taking into account user intent, the specific scientific domain, and their potential impact on the external environment. Then, we explore the underlying causes of these vulnerabilities and provide a scoping review of the limited existing works. Based on our analysis, we propose a triadic framework involving human regulation, agent alignment, and an understanding of environmental feedback (agent regulation) to mitigate these identified risks. Furthermore, we highlight the limitations and challenges associated with safeguarding AI scientists and advocate for the development of improved models, robust benchmarks, and comprehensive regulations.
1
A triadic mitigation framework is proposed: human regulation, agent alignment, and understanding environmental feedback (agent regulation).
2
AI scientists powered by large language models can autonomously conduct experiments and facilitate scientific discoveries across disciplines.
3
Risks depend on user intent, scientific domain, and potential impact on the external environment.
4
There are significant limitations and challenges to safeguarding AI scientists, necessitating better models, robust benchmarks, and comprehensive regulations.
5
These AI scientists introduce novel vulnerabilities and potential for misuse that require careful safety consideration.
6
Underlying causes of vulnerabilities are analyzed and existing literature on these vulnerabilities is limited.

AI scientists powered by large language models

Vulnerabilities and risks associated with autonomous AI scientists, including misuse, underlying causes, and mitigation via human regulation, agent alignment, and environmental feedback (safeguarding over autonomy)

Publication Details
Publication Date
2025-09-18
Journal
Publisher
ISSN
Cited by
32
Access Type
Author Information
Authors
Arman Cohan
Zhiyong Lu
Qiao Jin
Mark Gerstein
Zhuosheng Zhang
Meng Qu
Xiangru Tang
Wangchunshu Zhou
Tongxin Yuan
Kunlun Zhu
Yichi Zhang
Yilun Zhao
J. Tang
Dov Greenbaum
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%