跳到主内容
@wquguru
精选75Chubby♨️产品发布/更新

Accio开源CommerceAgentBench:107项电商任务实测Age

Most AI benchmarks test whether a model can give the right answer. CommerceAgent…

原文
发到 X

Most AI benchmarks test whether a model can give the right answer. CommerceAgentBench asks whether an agent can actually finish the work.

大多数AI基准测试检验的是模型能否给出正确答案,而CommerceAgentBench则考察智能体能否真正完成工作。

Accio has open-sourced 107 e-commerce tasks across procurement, product listings, operations, fulfillment and after-sales. Agents work across browsers, email, calendars, documents, APIs and files.

Accio已开源了涵盖采购、商品上架、运营、履约和售后等领域的107项电商任务。智能体需在浏览器、电子邮件、日历、文档、API和文件之间协同工作。

Crucially, they are not graded on what they claim to have done. The benchmark verifies what they actually changed, saved or submitted.

关键在于,它们并非依据声称完成的内容来评分。该基准验证的是它们实际更改、保存或提交的内容。

Take the Gmail procurement case. The agent must search roughly 300 messy emails, identify the real suppliers, reconstruct the latest quotes, compare six Incoterms and four currencies, calculate landed costs and detect payment fraud.

以Gmail采购案例为例,智能体必须搜索约300封杂乱邮件,识别真实供应商,重建最新报价,比较六种国际贸易术语和四种货币,计算到岸成本,并检测支付欺诈。

Then it must choose a supplier, label the relevant emails, save a reply draft and create a kickoff event.

随后,它必须选择供应商,标记相关邮件,保存回复草稿,并创建启动会议。

This is the kind of benchmark I find genuinely useful. It measures agents more like workers than chatbots. And the results show why human oversight still matters: the best observed run completed only 66 of 107 tasks, a 61.7% pass rate. And since 2026 is literally the year of agents, this is more important than ever.

我认为这类基准测试确实很有价值。它衡量智能体的方式更像是对待员工,而非聊天机器人。结果也表明为何人工监督仍然重要:最佳观察运行仅完成了107项任务中的66项,通过率为61.7%。而且,既然2026年确实是智能体之年,这一点比以往任何时候都更为关键。

Accio says the tasks draw on 10M SMB users, 1.6M conversations, 200K agent trajectories and Alibaba’s 27 years of e-commerce experience.

Accio表示,这些任务基于1000万中小企业用户、160万次对话、20万条智能体轨迹以及阿里巴巴27年的电商经验。

The project and task specifications are open source: https://github.com/Accio-org/CommerceAgentBench

项目及任务规范均为开源: https://github.com/Accio-org/CommerceAgentBench

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近