跳到主内容
@wquguru
精选72Rohan Paul论文研究

微软发布ThinkingBox:面向状态工作流的Agent可靠性沙盒基准

Very relevant Microsoft paper on agent reliability.

原文
发到 X

Very relevant Microsoft paper on agent reliability.

一篇关于智能体可靠性的非常相关的微软论文。

Succeeding once and being reliable are not the same thing.

成功一次与具备可靠性并非同一回事。

The best agent solved 91% of business tasks at least once but only 25% every time, so measure repeats.

表现最佳的智能体至少一次性解决了91%的业务任务,但每次都能解决的仅占25%,因此需要测量重复成功率。

The failures are also hard to spot from the outside.

这些失败从外部也很难察觉。

4 out of 5 failed runs ended politely and called a tool that writes to the database.

5次失败运行中有4次以礼貌的方式结束,并调用了写入数据库的工具。

The agent said the job was done, but the records said otherwise.

智能体声称工作已完成,但记录显示情况并非如此。

So Microsoft built ThinkingBox, a sandbox that runs an agent against real tools, a simulated customer, and a live backend, then checks the database afterward instead of reading the reply.

因此,微软构建了ThinkingBox,这是一个沙盒环境,让智能体在真实工具、模拟客户和实时后端上运行,随后检查数据库而非依赖回复内容。

– arxiv. org/abs/2608.19741

– arxiv.org/abs/2608.19741

Title: "One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows"

标题:《一次成功不等于可靠性:Thinkingbox,面向状态式业务流程中智能体的沙盒与基准测试》

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近