微软发布ThinkingBox:面向状态工作流的Agent可靠性沙盒基准
Very relevant Microsoft paper on agent reliability.
Very relevant Microsoft paper on agent reliability.
一篇关于智能体可靠性的非常相关的微软论文。
Succeeding once and being reliable are not the same thing.
成功一次与具备可靠性并非同一回事。
The best agent solved 91% of business tasks at least once but only 25% every time, so measure repeats.
表现最佳的智能体至少一次性解决了91%的业务任务,但每次都能解决的仅占25%,因此需要测量重复成功率。
The failures are also hard to spot from the outside.
这些失败从外部也很难察觉。
4 out of 5 failed runs ended politely and called a tool that writes to the database.
5次失败运行中有4次以礼貌的方式结束,并调用了写入数据库的工具。
The agent said the job was done, but the records said otherwise.
智能体声称工作已完成,但记录显示情况并非如此。
So Microsoft built ThinkingBox, a sandbox that runs an agent against real tools, a simulated customer, and a live backend, then checks the database afterward instead of reading the reply.
因此,微软构建了ThinkingBox,这是一个沙盒环境,让智能体在真实工具、模拟客户和实时后端上运行,随后检查数据库而非依赖回复内容。
– arxiv. org/abs/2608.19741
– arxiv.org/abs/2608.19741
Title: "One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows"
标题:《一次成功不等于可靠性:Thinkingbox,面向状态式业务流程中智能体的沙盒与基准测试》
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力