跳到主内容
@wquguru
精选75Rohan Paul论文研究

多轮长程规划物理:从预训练到后训练的单/多教师策略蒸馏

No amount of post-training cleanly fixes weak long-horizon foundations: noisy tr…

原文
发到 X

No amount of post-training cleanly fixes weak long-horizon foundations: noisy trajectories compound errors, sparse rewards misassign credit, and conflicting teachers trigger forgetting.

无论多少后期训练都无法彻底修复薄弱的长程基础:噪声轨迹会累积错误,稀疏奖励会错误分配信用,而相互冲突的教师模型会引发遗忘。

This paper finds, ff you want agents to stay reliable over long tasks, give them clean world-model and long-trajectory training first, use OPD (on-policy distillation) when reward signals get too sparse, and avoid merging teachers with incompatible planning strategies.

本文发现,如果你希望智能体在长任务中保持可靠,应首先提供清晰的世界模型和长轨迹训练,当奖励信号过于稀疏时使用OPD(同策略蒸馏),并避免合并具有不兼容规划策略的教师模型。

Suboptimal trajectories were especially damaging because small mistakes accumulated until middle and long tasks nearly collapsed.

次优轨迹尤其有害,因为小错误会累积,直到中长任务几乎崩溃。

For post-training, OPD handled longer, noisier settings better than outcome-reward GRPO because teacher feedback arrived throughout the trajectory instead of only at the end.

对于后期训练,OPD在处理更长、更嘈杂的环境时优于结果奖励GRPO,因为教师反馈在整个轨迹中提供,而不仅仅在结束时。

– arxiv. org/abs/2607.24720v1

– arxiv.org/abs/2607.24720v1

Title: "The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation"

标题:“多轮长程规划的物理学:从预训练到后期训练,通过单教师和多教师同策略智能体蒸馏”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近