Evaluation and Benchmarking of LLM Agents: A Survey
Оценка и бенчмаркинг агентов на основе больших языковых моделей: обзор
2025-08-03
SCID: 54.1/6qbz9z7p
Discuss with AI
LLM agent evaluationagent reliabilityagent safetybenchmarkingevaluation taxonomy
Figures from the paper
Abstract (AI)
The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area.This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a twodimensional taxonomy that organizes existing work along (1) evaluation objectives-what to evaluate, such as agent behavior, capabilities, reliability, and safety-and (2) evaluation process-how to evaluate, including interaction modes, datasets and benchmarks, metric computation methods, and tooling.In addition to taxonomy, we highlight enterprise-specific challenges, such as role-based access to data, the need for reliability guarantees, dynamic and longhorizon interactions, and compliance, which are often overlooked in current research.We also identify the future research directions, including holistic, more realistic, and scalable evaluation.This work aims to bring clarity to the fragmented landscape of agent evaluation and provide a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.
Key Findings
1
Enterprise evaluation must address role-based data access, reliability guarantees, dynamic and long-horizon interactions, and compliance requirements.
2
Future research should develop evaluation approaches that are holistic, realistic, and scalable for real-world LLM-agent deployment.
3
LLM-agent evaluation is complex and remains underdeveloped despite the rapid expansion of agent-based AI applications.
4
The evaluation-process dimension covers interaction modes, datasets and benchmarks, metric computation methods, and supporting tools.
5
The survey introduces a two-dimensional taxonomy organizing evaluation by objectives—behavior, capabilities, reliability, and safety—and by evaluation processes.
Research Object
LLM-based agents
Research Subject
evaluation and benchmarking of agent behavior, capabilities, reliability, and safety, including systematic assessment under realistic, dynamic, long-horizon, enterprise, and compliance constraints
Publication Details
Publication Date
2025-08-03
Journal
Publisher
ISSN
Cited by
57
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest