多轮长程规划物理:从预训练到后训练的单/多教师策略蒸馏
No amount of post-training cleanly fixes weak long-horizon foundations: noisy tr…
No amount of post-training cleanly fixes weak long-horizon foundations: noisy trajectories compound errors, sparse rewards misassign credit, and conflicting teachers trigger forgetting.
无论多少后期训练都无法彻底修复薄弱的长程基础:噪声轨迹会累积错误,稀疏奖励会错误分配信用,而相互冲突的教师模型会引发遗忘。
This paper finds, ff you want agents to stay reliable over long tasks, give them clean world-model and long-trajectory training first, use OPD (on-policy distillation) when reward signals get too sparse, and avoid merging teachers with incompatible planning strategies.
本文发现,如果你希望智能体在长任务中保持可靠,应首先提供清晰的世界模型和长轨迹训练,当奖励信号过于稀疏时使用OPD(同策略蒸馏),并避免合并具有不兼容规划策略的教师模型。
Suboptimal trajectories were especially damaging because small mistakes accumulated until middle and long tasks nearly collapsed.
次优轨迹尤其有害,因为小错误会累积,直到中长任务几乎崩溃。
For post-training, OPD handled longer, noisier settings better than outcome-reward GRPO because teacher feedback arrived throughout the trajectory instead of only at the end.
对于后期训练,OPD在处理更长、更嘈杂的环境时优于结果奖励GRPO,因为教师反馈在整个轨迹中提供,而不仅仅在结束时。
– arxiv. org/abs/2607.24720v1
– arxiv.org/abs/2607.24720v1
Title: "The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation"
标题:“多轮长程规划的物理学:从预训练到后期训练,通过单教师和多教师同策略智能体蒸馏”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力