NVIDIA 新论文:技能文档高分不代表运行时有效,应测轨迹增益
New Nvidia paper.
New Nvidia paper.
英伟达新论文。
A high-scoring skill document is not evidence that the skill helps the agent at runtime.
高分的技能文档并不能证明该技能在运行时对智能体有帮助。
Most of what an agent skill adds is discovery and workflow order, not a better final answer. So grade the trajectory, since final-answer scoring hides whether the agent ever found or followed the skill.
智能体技能增加的大部分内容是发现和工作流顺序,而非更好的最终答案。因此,应评估轨迹,因为最终答案评分掩盖了智能体是否找到或遵循了技能。
So run the same task twice, once with the skill loaded and once without, then compare.
所以,应运行同一任务两次,一次加载技能,一次不加载,然后进行比较。
NVIDIA's ACES holds task, model, harness, sandbox, grader, and supporting skills fixed, varies only whether the target skill is present, and reports the paired reward difference as Skill Lift.
英伟达的ACES系统固定任务、模型、框架、沙箱、评分器和辅助技能,仅变化目标技能是否存在,并将配对奖励差异报告为技能提升。
On production skills with both a document score and a live run, structural and judge scores correlate with measured lift at −0.0181 and −0.0266, indistinguishable from zero.
在既有文档评分又有实时运行的生产技能上,结构评分和裁判评分与实测提升的相关系数分别为−0.0181和−0.0266,与零无显著差异。
– arxiv. org/abs/2608.20614
– arxiv.org/abs/2608.20614
Title: "Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills"
标题:“评估技能,而非仅评估智能体:技能的智能体持续评估”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力