跳到主内容
精选86elvis论文研究

NVIDIA提出ACES框架:通过技能提升度评估Agent能力

Very interesting new paper from NVIDIA.

原文
推荐理由

Agent技能评估是落地痛点,这篇论文用严谨的对照实验推翻了传统的静态扫描标准,提出了可量化的动态评估方法,做Agent基建的同学值得参考。

Very interesting new paper from NVIDIA.

来自 NVIDIA 的一篇非常有趣的新论文。

(bookmark it)

(收藏它)

It takes a closer look at evaluating agent skills.

它深入探讨了如何评估智能体技能。

Enterprise teams are starting to leverage shared skill libraries, and the review gate is typically a scanner that checks structure, style, and security.

企业团队开始利用共享的技能库,而审查关卡通常是一个检查结构、风格和安全的扫描器。

NVIDIA measured whether that gate predicts anything.

NVIDIA 测量了该关卡是否能预测任何内容。

Across 145 real skills from internal and public catalogs, structural scan scores correlate with LLM-judge quality at a Spearman rho of 0.14.

在内部和公共目录中的 145 个真实技能上,结构扫描分数与 LLM 裁判质量的相关性(Spearman rho)为 0.14。

ACES proposes Skill Lift instead.

ACES 提出了“技能提升”(Skill Lift)方法。

In other words, run the same task twice under the same model, sandbox, workspace, and scorer, once with the skill loaded and once without. Then you measure the difference in what the agent completed.

换句话说,在相同的模型、沙箱、工作区和评分器下运行两次相同任务,一次加载技能,一次不加载。然后你衡量智能体完成情况的差异。

They scored 947 paired cases from 58 production skills across four harnesses, normalizing trajectories into a shared Agent Trajectory Interchange Format, so results compare across harnesses.

他们对来自四个框架的 58 个生产技能中的 947 对案例进行了评分,将轨迹标准化为共享的智能体轨迹交换格式,以便在不同框架之间比较结果。

They fins that the largest process-metric gains appear in skill execution, behavior check, and skill efficiency.

他们发现最大的过程指标增益出现在技能执行、行为检查和技能效率方面。

Paper: https://arxiv.org/abs/2608.20614

论文:https://arxiv.org/abs/2608.20614

Track more trending AI papers in our academy: https://academy.dair.ai/

在我们的学院中追踪更多热门 AI 论文:https://academy.dair.ai/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近