Skip to main content
Aggregate AI 摘要 arXiv cs.AI 人工智能 7 Sep 2026 - 14:00

$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction

RSS 官方收录 · 可信分层展示

关键摘要

τ^τ-Bench首次评测AI编码代理构建真实客服代理的能力,通过率仅23.9%

  • benchmark要求AI编码代理基于真实业务数据、客户需求数和生产API端到端构建客服代理
  • 评估覆盖4个领域53项任务,最强配置(Claude Opus 5+Claude Code)通过率23.9%
  • 专家人工方案通过率82.2%,AI主要失败于浅层查询、缺乏客户沟通及架构实验不足

AI 摘要 · 来源可核验

正文提要

arXiv:2609.04611v1 Announce Type: new Abstract: LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a real client engagement. We introduce $\tau^\tau$-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task. A developer agent is given the records a business actually keeps, a client who holds requirements, a production API that operations must run through, a codebase to inherit, and limits on serving cost and models: the same starting point a real engagement provides. From these it must deliver a complete customer-service agent, scored by deploying that agent against held-out simulated users. Across 53 tasks spanning four domains, the strongest configuration, Claude Opus 5 under Claude Code, passes just 23.9% of evaluation simulations. Meanwhile, an expert-authored reference ceiling scores 82.2%. The failures mirror ones human agent developers see: models issue shallow queries in place of deep comprehension of the records, communicate almost nothing to the client, and experiment too little with agent architecture and serving spend, shipping the first design that runs. We aim for $\tau^\tau$-bench to turn the work of cooperative agent building into a measurable target for coding agents.

来源:https://arxiv.org/abs/2609.04611

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表