微软新研究:坏技能库会让Agent失败率上升
Very interesting new paper from Microsoft and colleagues.
Very interesting new paper from Microsoft and colleagues.
微软及其同事发表了一篇非常有趣的新论文。
(bookmark it)
(收藏它)
Skill libraries are used in every major harness on the assumption that more guidance is free. This work measures what a bad skill actually costs you.
技能库被用于所有主要框架中,其假设是更多的指导是免费的。这项工作衡量了一个糟糕的技能实际上会给你带来多少成本。
They attribute 307 agent failures to specific loaded skills, 125 functional failures and 182 efficiency regressions, by comparing each skill-guided run against a matched reference run that solves the same task.
他们通过将每次技能引导的运行与解决相同任务的匹配参考运行进行比较,将307个代理失败归因于特定的加载技能,其中125个功能失败和182个效率回归。
The failures rarely come from irrelevant skills. Seemingly relevant skills push the agent to incorrectly implement or omit something the task required.
失败很少来自不相关的技能。看似相关的技能会促使代理错误地实现或遗漏任务要求的内容。
Cost regressions are not explained by prompt length either. The largest source is excessive verification at 67 cases, followed by heavy implementation pipelines at 30 cases. It turns out that skills quietly turn validation checklists into mandatory work.
成本回归也不能用提示长度来解释。最大的来源是过度验证,有67例,其次是繁重的实现流程,有30例。事实证明,技能会悄悄地将验证检查表变成强制性工作。
Paper: https://arxiv.org/abs/2608.11888
论文:https://arxiv.org/abs/2608.11888
Track more trending AI papers in our academy: https://academy.dair.ai/
在我们的学院中追踪更多热门AI论文:https://academy.dair.ai/
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力