ττ-bench评测:顶级编程Agent仅23.9%通过率
The best coding-agent setup passed only 23.9% of held-out customer simulations o…
用硬核数据打破“代码强即Agent强”的迷思,揭示了需求理解与交互的关键短板,对Agent研发极具参考价值。
The best coding-agent setup passed only 23.9% of held-out customer simulations on this benchmark.
在此基准测试中,表现最佳的编码智能体配置仅通过了23.9%的保留客户模拟用例。
ττ-bench finds that today's coding agents fail most realistic agent-building work because they do not understand requirements deeply enough, so improving code generation alone will not solve the problem.
ττ-bench 发现,当前的编码智能体在大多数真实的智能体构建工作中失败,因为它们对需求的理解不够深入,因此仅改进代码生成无法解决问题。
ττ-bench treats agent building like a real client job.
ττ-bench 将智能体构建视为真实的客户项目。
The model gets scattered company records, a client, an API, existing code, and a budget, then must ship an agent for unseen customer requests.
模型会获得分散的公司记录、一位客户、一个 API、现有代码以及预算,然后必须为未见的客户需求交付一个智能体。
The best setup, Claude Opus 5 with Claude Code, passed just 23.9% of held-out customer simulations across 53 tasks, versus 82.2% for the expert-authored reference.
最佳配置(Claude Opus 5 搭配 Claude Code)在53项任务中仅通过了23.9%的保留客户模拟用例,而专家编写的参考实现达到了82.2%。
Most failures were outside pure coding.
大多数失败案例并非源于纯编码问题。
Agents searched records instead of understanding them, barely questioned the client, reused familiar designs, and wrote tests that often missed their own errors.
智能体只是搜索记录而非真正理解它们,几乎不向客户提问,复用熟悉的设计模式,且编写的测试往往未能发现自身的错误。
On client-enabled tasks, builds asking 0 questions averaged 0.16; those asking 4+ averaged 0.50.
在启用客户交互的任务中,提出0个问题的构建平均得分为0.16;提出4个及以上问题的构建平均得分为0.50。
So, better coding models are not enough, we also need coding agents that gather requirements, compare designs, use budgets intelligently, and run tests that expose their own blind spots before deployment.
因此,仅靠更强大的编码模型是不够的,我们还需要能够收集需求、比较设计方案、智能使用预算,并在部署前运行能暴露自身盲点的测试的编码智能体。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力