跳到主内容
精选75elvis论文研究

论文揭示智能体后训练策略锁定现象与改进尝试

Great paper if you are tracking progress in recursive self-improvement (RSI).

原文

Great paper if you are tracking progress in recursive self-improvement (RSI).

(bookmark it)

There is so much hype around RSI, so I think it's worth understanding why current models are not able to do this properly yet.

Issues range from "lack of creativity" of models to getting stuck in a local optimum.

This work tries to provide more insights into whether agents can really post-train other agents.

Here is the most interesting finding reported in the paper: "the agent’s training strategy is locked in at the very beginning, and the entire remaining budget is spent on local adjustments within the selected strategy."

They analyzed a large corpus of publicly released post-training trajectories. Across tasks, the agent locks in its training strategy at the very first step and spends the entire remaining budget on local adjustments inside it.

They then tried three escalating fixes. An experience-driven scaffold lifted execution broadly, worth 12.6 points on GSM8K and 40.8 on HumanEval, and the strategy stayed frozen.

Human guidance redirected the opening choice, and the agent slid back into local loops once training began. Extra inference compute paid off on easy tasks and did almost nothing on the hardest one.

What agents lack here is a way to reconsider strategy while execution is still running.

Paper: https://arxiv.org/abs/2608.19072

Track more trending AI papers in our academy: https://academy.dair.ai/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近