Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

На пути к заслуживающему доверия агентному искусственному интеллекту: всесторонний обзор безопасности, робастности, конфиденциальности и системной защиты
Jinhu Qi, Muzhi Li, Jiahong Liu, Yuqin Shu, D S F Yu, Shicheng Ma, Wenqian Cui, Yiyang Zhao, Yuehua Chen, Ruoxi Jiang, Irwin King, Zenglin Xu
2026-04-29

large language modelsprivacy and system securityruntime monitoring and verificationsafety and robustnesstrustworthy agentic AI
Agentic AI systems—Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions—can execute complex tasks autonomously, but their multi-step trajectories introduce new failure modes that challenge trustworthiness. This survey provides a focused examination of trustworthy agentic AI through two core dimensions that are critical for high-risk deployments: Safety and Robustness and Privacy and System Security. For each dimension, we clarify key concepts, identify where risks emerge along the agent workflow, and summarize stage-targeted mitigation strategies. Other trustworthiness aspects (value alignment, transparency, fairness, and accountability) are discussed as relevant context rather than parallel chapters. To support consistent comparison and deployment decisions, we consolidate evaluation into a unified metrics-and-benchmarks hub, emphasizing both outcome and process signals (e.g., constraint violations, trace completeness, and adversarial success rates) and offering scenario-to-metric guidance for release gating. We conclude by outlining open challenges such as self-evolving agents, runtime monitoring and verification, privacy-preserving personalization, and the trust–utility trade-off, and present a case study of real-world security failures in open-source agentic systems (OpenClaw/Moltbook). Our goal is to serve as a practical reference for researchers and practitioners building trustworthy agentic systems in high-stakes environments.
1
A unified metrics-and-benchmarks framework combines outcome and process signals, including constraint violations, trace completeness, and adversarial success rates, to guide release gating.
2
Agentic AI’s planning, tool use, memory, and long-horizon interactions create distinctive multi-step failure modes that complicate trustworthy deployment.
3
It maps risks across agent workflows and summarizes mitigation strategies targeted to specific stages of agent operation.
4
Major open challenges include self-evolving agents, runtime monitoring and verification, privacy-preserving personalization, trust–utility trade-offs, and real-world security failures in open-source systems.
5
The survey organizes trustworthy agentic AI around two high-risk dimensions: safety and robustness, and privacy and system security.

agentic AI systems based on large language models with planning, tool use, memory, and long-horizon interactions

trustworthiness, focusing on safety, robustness, privacy, and system security across autonomous multi-step workflows

Publication Details
Publication Date
2026-04-29
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Jinhu Qi
Muzhi Li
Jiahong Liu
Yuqin Shu
D S F Yu
Shicheng Ma
Wenqian Cui
Yiyang Zhao
Yuehua Chen
Ruoxi Jiang
Irwin King
Zenglin Xu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%