微软AutoSaddler:自动优化Agent harness
Impressive new paper from Microsoft and colleagues.
做Agent工程的同学必看,AutoSaddler把harness调优从手工变成自动化闭环,三个基准都涨了约10分,值得拿你的失败轨迹试试这套离线修补思路。
Impressive new paper from Microsoft and colleagues.
微软及其同事发表了一篇令人印象深刻的新论文。
Harness design is still hand-tuned almost everywhere. This work present an automated loop to optimize the harness.
几乎在所有地方,装备设计仍然依赖手工调整。这项工作提出了一个自动化循环来优化装备。
They introduce AutoSaddler, which treats the agent harness as code and learns to patch it offline from failure traces.
他们引入了AutoSaddler,将代理装备视为代码,并学习从失败轨迹中离线修补它。
It runs mini batches of tasks, diagnoses what broke, generates structured patches to prompts, tool configurations, and control logic, then keeps an update only if it survives validation.
它运行小批量任务,诊断出错原因,生成针对提示、工具配置和控制逻辑的结构化补丁,并且只有在通过验证后才保留更新。
Gains of 9.0 points on GAIA2, 9.6 on SWE-Bench Pro, and 10.0 on Terminal-Bench 2.0 over the corresponding base harnesses.
在GAIA2上提升了9.0分,在SWE-Bench Pro上提升了9.6分,在Terminal-Bench 2.0上提升了10.0分,均优于相应的基础装备。
Deep debugging beats shallow reflection, targeted edits beat unconstrained editing, and generalization-aware selection beats repairing the one trajectory in front of you.
深度调试优于浅层反思,针对性编辑优于无约束编辑,泛化感知选择优于修复眼前的一条轨迹。
Paper: https://arxiv.org/abs/2608.23041
论文:https://arxiv.org/abs/2608.23041
Track more trending AI papers in our academy: https://academy.dair.ai/
在我们的学院中追踪更多热门AI论文:https://academy.dair.ai/
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力