CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation in Real-World API Interactions
CToolEval: китайский бенчмарк для оценки агентов на основе больших языковых моделей при взаимодействии с API в реальных условиях
2024-01-01
SCID: 54.1/7m3tsnd3
Discuss with AI
API interactionsCToolEval benchmarkChinese societal applicationsLLM-powered agentstool invocation
Figures from the paper
Abstract (AI)
Assessing the capabilities of large language models (LLMs) as agents in decision making and operational tasks is crucial for the development of LLM-as-agent service.We propose CToolEval, a benchmark designed to evaluate LLMs in the context of Chinese societal applications, featuring 398 APIs across 27 widelyused Apps (e.g., Apps for shopping, map, music, travel, etc.) that cover 14 domains.We further present an evaluation framework that simulates real-life scenarios, to facilitate the assessment of tool invocation ability of LLMs for tool learning and task completion ability for user interation.Our extensive experiments with CToolEval evaluate 11 LLMs, revealing that while GPT-3.5-turboexcels in tool invocation, Chinese LLMs usually struggle with issues like hallucination and a lack of comprehensive tool understanding.Our findings highlight the need for further refinement in decision-making capabilities of LLMs, offering insights into bridging the gap between current functionalities and agent-level performance.To promote further research for LLMs to fully act as reliable agents in complex, real-world situations, we release our data and codes at https: //github.com/tjunlp-lab/CToolEval.
Key Findings
1
CToolEval introduces a Chinese benchmark with 398 APIs from 27 widely used applications spanning 14 societal domains.
2
Chinese LLMs commonly exhibit hallucinations and insufficient comprehensive understanding of available tools.
3
Experiments across 11 LLMs show that GPT-3.5-turbo performs best in tool invocation among the evaluated models.
4
The benchmark framework simulates real-life scenarios to evaluate both LLM tool invocation for tool learning and task completion during user interactions.
5
The results indicate that current LLMs require improved decision-making capabilities to achieve reliable agent-level performance in complex real-world settings.
Research Object
LLM-powered agents interacting with APIs in Chinese societal applications
Research Subject
Tool invocation and task-completion capabilities, including hallucination and comprehensive tool understanding, evaluated in simulated real-life scenarios
Publication Details
Publication Date
2024-01-01
Journal
Publisher
ISSN
Cited by
6
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest