跳到主内容
@wquguru
精选75elvis论文研究

Harness-IF:用执行证据评估编码Agent规则有效性

If you write rules in an AGENTS.md, this one is worth your time.

原文
发到 X

If you write rules in an AGENTS.md, this one is worth your time.

如果你在AGENTS.md中编写规则,那么这条规则值得你花时间阅读。

When a coding agent follows your rule, it may have been going to do that anyway.

当编码代理遵循你的规则时,它可能本来就会这样做。

Harness-IF separates the two by scoring 256 rules one at a time from execution evidence, then re-running every task with the rule withheld across nine probe builds to find which rules actually oppose the model's defaults.

Harness-IF通过从执行证据中逐一评分256条规则,然后在九个探针构建中重新运行每个任务并省略该规则,以找出哪些规则实际上与模型的默认行为相悖,从而将两者区分开来。

Across 12 frontier models, raw accuracy runs 72.1 to 85.9%, and Against-Prior Accuracy runs 66.1 to 78.6%. Every model gets worse once coincidence is stripped out, by 3.6 to 7.4 points.

在12个前沿模型中,原始准确率在72.1%到85.9%之间,而对抗先验准确率在66.1%到78.6%之间。一旦去除巧合因素,每个模型的性能都会下降3.6到7.4个百分点。

One finding worth flagging. Precedence does not follow prompt depth. System prompts, project files, and user instructions all outrank tool and skill descriptions.

一个值得注意的发现是:优先级并不遵循提示深度。系统提示、项目文件和用户指令的优先级都高于工具和技能描述。

Paper: https://arxiv.org/abs/2608.11727

论文:https://arxiv.org/abs/2608.11727

Track more trending AI papers in our academy: https://academy.dair.ai/

在我们的学院中追踪更多热门AI论文:https://academy.dair.ai/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近