Meta新研究:智能体离线学习工具编排策略
New research from Meta.
New research from Meta.
Meta的新研究。
Agent harnesses are still mostly authored by hand.
代理工具(Agent harnesses)大多仍由手工编写。
This makes it hard to tune robust agent harnesses for long-horizon tasks.
这使得为长时程任务调整健壮的代理工具变得困难。
In this new work, agents learn harness policies offline and deploy them to construct and update external harness state online during runtime task execution.
在这项新工作中,代理离线学习工具策略,并在在线任务执行期间部署该策略来构建和更新外部工具状态。
EvoHarness-RL learns that policy instead. Belief, Progress, and Experience are exposed as harness state the policy can act on.
EvoHarness-RL学习该策略。信念、进度和经验作为工具状态暴露给策略,供其操作。
Supervised harness fine-tuning teaches the action space, then cost-aware GRPO explores when to read, update, and consolidate during a long run. Qwen3-8B reaches 96.9% on ALFWorld.
监督式工具微调教授动作空间,然后成本感知的GRPO在长时间运行中探索何时读取、更新和整合。Qwen3-8B在ALFWorld上达到96.9%的准确率。
Two dynamics come out of the training.
训练中出现了两种动态。
> Harness annealing means recurring harness-use patterns get absorbed into the model policy, and the agent shifts from frequent calls toward selective access.
> 工具退火意味着重复出现的工具使用模式被吸收到模型策略中,代理从频繁调用转向选择性访问。
> Harness evolution means progress updates and experience consolidation compress the workspace into a compact task-adaptive state.
> 工具进化意味着进度更新和经验整合将工作空间压缩为紧凑的任务自适应状态。
This shows that long-horizon agents get more from a trainable coordination policy than from bigger tools or larger memories.
这表明,对于长时程代理,可训练的协调策略比更大的工具或更大的记忆更有价值。
Paper: https://arxiv.org/abs/2608.05446
论文:https://arxiv.org/abs/2608.05446
Track more trending AI papers in our academy: https://academy.dair.ai/
在我们的学院中追踪更多热门AI论文:https://academy.dair.ai/
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力