新论文:智能体能否真正后训练其他智能体?
Finally a good paper testing whether agents can really post-train other agents.
Finally a good paper testing whether agents can really post-train other agents.
终于有一篇好论文测试了智能体是否真的能对其他智能体进行后训练。
(bookmark it)
(收藏它)
They analyzed a large corpus of publicly released post-training trajectories. Across tasks, the agent locks in its training strategy at the very first step and spends the entire remaining budget on local adjustments inside it.
他们分析了一个大型公开的后训练轨迹语料库。跨任务来看,智能体在第一步就锁定了其训练策略,并将剩余的整个预算花在其中的局部调整上。
They then tried three escalating fixes. An experience-driven scaffold lifted execution broadly, worth 12.6 points on GSM8K and 40.8 on HumanEval, and the strategy stayed frozen.
然后他们尝试了三种逐步升级的修复方法。一种经验驱动的脚手架广泛提升了执行能力,在GSM8K上提升了12.6分,在HumanEval上提升了40.8分,但策略仍然保持冻结。
Human guidance redirected the opening choice, and the agent slid back into local loops once training began. Extra inference compute paid off on easy tasks and did almost nothing on the hardest one.
人类指导重定向了初始选择,但一旦训练开始,智能体又滑回了局部循环。额外的推理计算在简单任务上效果显著,而在最难的任务上几乎毫无作用。
What agents lack here is a way to reconsider strategy while execution is still running.
这里智能体所缺乏的是一种在执行过程中重新考虑策略的方式。
Paper: https://arxiv.org/abs/2608.19072
论文:https://arxiv.org/abs/2608.19072
Track more trending AI papers in our academy: https://academy.dair.ai/
在我们的学院中追踪更多热门AI论文:https://academy.dair.ai/
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力