论文:LLM Agent技能积累未必带来性能提升,上下文反馈更关键
Agents can accumulate hundreds of skills without becoming proportionally better,…
直接挑战了当前Agent开发中“技能库”有效性的假设,用数据证明上下文记忆往往比显式技能抽象更可靠,对Agent架构设计有重要参考价值。
Agents can accumulate hundreds of skills without becoming proportionally better, which makes skill consolidation and reuse a bigger problem than simply generating more skills.
智能体可以积累数百项技能,但并不会因此按比例变得更强,这使得技能的整合与复用成为一个比单纯生成更多技能更严峻的问题。
If you want an agent to improve over repeated work, keeping its context and feedback is already useful; autonomous skill creation is still unreliable except where reusable procedures really matter.
如果你希望智能体在重复工作中实现改进,保留其上下文和反馈已经很有用;除非可复用的流程确实至关重要,否则自主创建技能仍然不可靠。
Agents can get better from experience, but this paper finds that explicit skill libraries are not yet consistently better than carrying forward context and feedback.
智能体可以通过经验获得提升,但这篇论文发现,明确的技能库并不总是比延续上下文和反馈表现更好。
ContinualSkillBench gives agents 100 connected tasks in each of 5 domains, lets them keep feedback and update reusable skills, and compares that with solving every task from scratch.
ContinualSkillBench 为智能体在每个领域的 100 个关联任务中提供环境,允许它们保留反馈并更新可复用技能,并将其与从头解决每个任务的情况进行对比。
Sequential execution improved normalized reward in 14 of 15 model-domain settings, a 16.9% relative gain overall.
在 15 个模型-领域设置中的 14 个里,顺序执行提升了归一化奖励,整体相对增益为 16.9%。
But the ablation changes the takeaway: on GPT-5.3-Codex across Law, Finance, and Healthcare, pure in-context learning averaged 0.605 normalized reward versus 0.602 with explicit skill maintenance.
但消融实验改变了结论:在 GPT-5.3-Codex 上,针对法律、金融和医疗领域,纯上下文学习的平均归一化奖励为 0.605,而带有明确技能维护的为 0.602。
So much of the gain appears to come from carrying forward context and feedback, not from agents reliably abstracting reusable skills.
因此,大部分收益似乎来自于延续上下文和反馈,而非智能体可靠地抽象出可复用技能。
– arxiv. org/abs/2608.03874
– arxiv.org/abs/2608.03874
Title: "ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?"
标题:《ContinualSkillBench:LLM 智能体真的能够进化其能力吗?》
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力