CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation in Real-World API Interactions

CToolEval: китайский бенчмарк для оценки агентов на основе больших языковых моделей при взаимодействии с API в реальных условиях
Zishan Guo, Yufei Huang, Deyi Xiong
2024-01-01

API interactionsCToolEval benchmarkChinese societal applicationsLLM-powered agentstool invocation
Assessing the capabilities of large language models (LLMs) as agents in decision making and operational tasks is crucial for the development of LLM-as-agent service.We propose CToolEval, a benchmark designed to evaluate LLMs in the context of Chinese societal applications, featuring 398 APIs across 27 widelyused Apps (e.g., Apps for shopping, map, music, travel, etc.) that cover 14 domains.We further present an evaluation framework that simulates real-life scenarios, to facilitate the assessment of tool invocation ability of LLMs for tool learning and task completion ability for user interation.Our extensive experiments with CToolEval evaluate 11 LLMs, revealing that while GPT-3.5-turboexcels in tool invocation, Chinese LLMs usually struggle with issues like hallucination and a lack of comprehensive tool understanding.Our findings highlight the need for further refinement in decision-making capabilities of LLMs, offering insights into bridging the gap between current functionalities and agent-level performance.To promote further research for LLMs to fully act as reliable agents in complex, real-world situations, we release our data and codes at https: //github.com/tjunlp-lab/CToolEval.
1
CToolEval introduces a Chinese benchmark with 398 APIs from 27 widely used applications spanning 14 societal domains.
2
Chinese LLMs commonly exhibit hallucinations and insufficient comprehensive understanding of available tools.
3
Experiments across 11 LLMs show that GPT-3.5-turbo performs best in tool invocation among the evaluated models.
4
The benchmark framework simulates real-life scenarios to evaluate both LLM tool invocation for tool learning and task completion during user interactions.
5
The results indicate that current LLMs require improved decision-making capabilities to achieve reliable agent-level performance in complex real-world settings.

LLM-powered agents interacting with APIs in Chinese societal applications

Tool invocation and task-completion capabilities, including hallucination and comprehensive tool understanding, evaluated in simulated real-life scenarios

Publication Details
Publication Date
2024-01-01
Journal
Publisher
ISSN
Cited by
6
Access Type
Author Information
Authors
Zishan Guo
Yufei Huang
Deyi Xiong
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%