← Back to Paper List

CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation in Real-World API Interactions

Zishan Guo, Yufei Huang, Deyi Xiong
Tianjin University, Tianjin, China
Annual Meeting of the Association for Computational Linguistics (2024)
Agent Benchmark

📝 Paper Summary

Benchmark datasets Multi-call tool use with flexible plan
CToolEval is a fine-grained benchmark with 398 real-world APIs to evaluate the tool invocation and task completion capabilities of LLM-powered agents in Chinese societal applications.
Core Problem
Existing tool evaluation benchmarks often rely on advanced models like GPT-4 for scoring, which introduces bias and fails to accurately evaluate dynamic, real-time answers.
Why it matters:
  • Current benchmarks struggle to assess dynamic real-time answers (e.g., weather forecasts) where the ground truth changes.
  • Relying on LLM-as-a-judge (like GPT-4) introduces subjective bias and scalability issues.
  • A lack of rigorous Chinese-specific tool-learning benchmarks limits the development of localized LLM agents.
Concrete Example: For a real-time query like 'highest temperature tomorrow', neither humans nor GPT-4 can judge correct answers in advance; CToolEval solves this by dynamically extracting real-time answers from API observations.
Key Novelty
CToolEval Benchmark and Dynamic Evaluation Framework
  • Categorizes interactions into fixed-answer, open-ended, real-time, and operational types to apply tailored, accuracy-based metrics.
  • Evaluates real-time queries dynamically by storing API links and parameters to fetch live reference answers, bypassing LLM-as-a-judge limitations.
  • Implements a rollback function for operational queries to reverse database changes after successful API calls, ensuring repeatable evaluation.
Architecture
Architecture Figure Figure 1
The framework of CToolEval covering interactions in real-world scenarios across 27 Apps and the evaluation of LLMs' tool invocation and task completion capabilities.
Evaluation Highlights
  • GPT-3.5-turbo achieves a 61.94% overall tool matching rate, significantly outperforming the best Chinese open-source model (InternLM-chat-20B at 21.02%).
  • In task completion, GPT-3.5-turbo scores 45.45% overall, while InternLM-chat-20B achieves only 15.20%.
  • Fine-tuning Qwen-7B-chat with QLoRA on CToolEval improves its overall tool matching rate from 0.00% to 63.60%, demonstrating the dataset's utility for training.
Breakthrough Assessment
7/10
Provides a rigorous, real-world benchmark for Chinese LLMs with a novel dynamic evaluation for real-time queries, though limited by API volatility.
×