Evaluation and Benchmarking of LLM Agents: A Survey

Оценка и бенчмаркинг агентов на основе больших языковых моделей: обзор
Mahmoud Mohammadi, Yipeng Li, Jane C. Lo, Wendy Yip
2025-08-03

LLM agent evaluationagent reliabilityagent safetybenchmarkingevaluation taxonomy
The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area.This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a twodimensional taxonomy that organizes existing work along (1) evaluation objectives-what to evaluate, such as agent behavior, capabilities, reliability, and safety-and (2) evaluation process-how to evaluate, including interaction modes, datasets and benchmarks, metric computation methods, and tooling.In addition to taxonomy, we highlight enterprise-specific challenges, such as role-based access to data, the need for reliability guarantees, dynamic and longhorizon interactions, and compliance, which are often overlooked in current research.We also identify the future research directions, including holistic, more realistic, and scalable evaluation.This work aims to bring clarity to the fragmented landscape of agent evaluation and provide a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.
1
Enterprise evaluation must address role-based data access, reliability guarantees, dynamic and long-horizon interactions, and compliance requirements.
2
Future research should develop evaluation approaches that are holistic, realistic, and scalable for real-world LLM-agent deployment.
3
LLM-agent evaluation is complex and remains underdeveloped despite the rapid expansion of agent-based AI applications.
4
The evaluation-process dimension covers interaction modes, datasets and benchmarks, metric computation methods, and supporting tools.
5
The survey introduces a two-dimensional taxonomy organizing evaluation by objectives—behavior, capabilities, reliability, and safety—and by evaluation processes.

LLM-based agents

evaluation and benchmarking of agent behavior, capabilities, reliability, and safety, including systematic assessment under realistic, dynamic, long-horizon, enterprise, and compliance constraints

Publication Details
Publication Date
2025-08-03
Journal
Publisher
ISSN
Cited by
57
Access Type
Author Information
Authors
Mahmoud Mohammadi
Yipeng Li
Jane C. Lo
Wendy Yip
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%