跳到主内容
@wquguru
精选75Rohan Paul论文研究

AutoSaddler:自动优化Agent框架,仅保留泛化补丁

Automatically patching an agent's harness is easy; keeping the patches that help…

原文
发到 X

Automatically patching an agent's harness is easy; keeping the patches that help is the hard part.

自动修补代理的“缰绳”(harness)很容易;难的是保留那些真正有帮助的修补。

So if you let a model rewrite your prompts, tools, and hooks from failure traces, spend the rollout budget on a held-out split that catches collateral damage.

因此,如果你让模型根据失败轨迹重写提示词、工具和钩子,应将部署预算花在保留的验证集上,以捕捉附带损害。

Harness optimization does work, but this paper finds that without a check on whether an update generalizes, the result lands below the hand-written harness it started from.

缰绳优化确实有效,但本文发现,如果不检查更新是否具有泛化性,结果会低于其最初的手写缰绳。

AutoSaddler diagnoses failed traces from a mini-batch, treats the harness as code, patches prompts, tools, and middleware, and keeps only updates that also improve a held-out development set.

AutoSaddler 从小批量数据中诊断失败轨迹,将缰绳视为代码,修补提示词、工具和中间件,并且只保留那些也能改进保留开发集的更新。

If you're tuning an agent by hand or with an LLM in the loop, hold out a set of tasks that the patch was not written for, and score fixes minus regressions rather than fixes alone.

如果你手动或通过 LLM 在循环中调整代理,请保留一组修补未针对的任务,并评分修复减去回退,而不仅仅是修复本身。

– arxiv. org/abs/2608.23041

– arxiv.org/abs/2608.23041

Title: "AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces"

标题:“AutoSaddler:通过代理执行轨迹进行持久更新的自动缰绳优化”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近