跳到主内容
精选85elvis论文研究

微软AutoSaddler:自动优化Agent harness

Impressive new paper from Microsoft and colleagues.

原文
推荐理由

做Agent工程的同学必看,AutoSaddler把harness调优从手工变成自动化闭环,三个基准都涨了约10分,值得拿你的失败轨迹试试这套离线修补思路。

Impressive new paper from Microsoft and colleagues.

微软及其同事发表了一篇令人印象深刻的新论文。

Harness design is still hand-tuned almost everywhere. This work present an automated loop to optimize the harness.

几乎在所有地方,装备设计仍然依赖手工调整。这项工作提出了一个自动化循环来优化装备。

They introduce AutoSaddler, which treats the agent harness as code and learns to patch it offline from failure traces.

他们引入了AutoSaddler,将代理装备视为代码,并学习从失败轨迹中离线修补它。

It runs mini batches of tasks, diagnoses what broke, generates structured patches to prompts, tool configurations, and control logic, then keeps an update only if it survives validation.

它运行小批量任务,诊断出错原因,生成针对提示、工具配置和控制逻辑的结构化补丁,并且只有在通过验证后才保留更新。

Gains of 9.0 points on GAIA2, 9.6 on SWE-Bench Pro, and 10.0 on Terminal-Bench 2.0 over the corresponding base harnesses.

在GAIA2上提升了9.0分,在SWE-Bench Pro上提升了9.6分,在Terminal-Bench 2.0上提升了10.0分,均优于相应的基础装备。

Deep debugging beats shallow reflection, targeted edits beat unconstrained editing, and generalization-aware selection beats repairing the one trajectory in front of you.

深度调试优于浅层反思,针对性编辑优于无约束编辑,泛化感知选择优于修复眼前的一条轨迹。

Paper: https://arxiv.org/abs/2608.23041

论文:https://arxiv.org/abs/2608.23041

Track more trending AI papers in our academy: https://academy.dair.ai/

在我们的学院中追踪更多热门AI论文:https://academy.dair.ai/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近