AutoSaddler:自动优化Agent框架,仅保留泛化补丁
Automatically patching an agent's harness is easy; keeping the patches that help…
Automatically patching an agent's harness is easy; keeping the patches that help is the hard part.
自动修补代理的“缰绳”(harness)很容易;难的是保留那些真正有帮助的修补。
So if you let a model rewrite your prompts, tools, and hooks from failure traces, spend the rollout budget on a held-out split that catches collateral damage.
因此,如果你让模型根据失败轨迹重写提示词、工具和钩子,应将部署预算花在保留的验证集上,以捕捉附带损害。
Harness optimization does work, but this paper finds that without a check on whether an update generalizes, the result lands below the hand-written harness it started from.
缰绳优化确实有效,但本文发现,如果不检查更新是否具有泛化性,结果会低于其最初的手写缰绳。
AutoSaddler diagnoses failed traces from a mini-batch, treats the harness as code, patches prompts, tools, and middleware, and keeps only updates that also improve a held-out development set.
AutoSaddler 从小批量数据中诊断失败轨迹,将缰绳视为代码,修补提示词、工具和中间件,并且只保留那些也能改进保留开发集的更新。
If you're tuning an agent by hand or with an LLM in the loop, hold out a set of tasks that the patch was not written for, and score fixes minus regressions rather than fixes alone.
如果你手动或通过 LLM 在循环中调整代理,请保留一组修补未针对的任务,并评分修复减去回退,而不仅仅是修复本身。
– arxiv. org/abs/2608.23041
– arxiv.org/abs/2608.23041
Title: "AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces"
标题:“AutoSaddler:通过代理执行轨迹进行持久更新的自动缰绳优化”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力