字节跳动新论文:测试智能体是否真正遵循指令
New ByteDance paper asks a much better question about AI agents: did the instruc…
New ByteDance paper asks a much better question about AI agents: did the instruction actually change anything?
字节跳动新论文提出了一个关于AI代理的更好的问题:指令真的改变了什么吗?
We might be overestimating how controllable AI agents are because our tests often agree with their defaults.
我们可能高估了AI代理的可控性,因为我们的测试往往与其默认行为一致。
Want to know whether your agent actually follows instructions? Give it a rule that goes against what it normally does.
想知道你的代理是否真正遵循指令?给它一条与其常规行为相悖的规则试试。
Harness-IF tests that harder case: rules that push against a model's default behavior, spread across the places coding agents actually read, including system prompts, tool descriptions, skill descriptions, project files, and user instructions.
Harness-IF测试了更困难的情况:推动模型偏离默认行为的规则,分布在编码代理实际阅读的各个地方,包括系统提示、工具描述、技能描述、项目文件和用户指令。
Across 12 frontier models and 60 multi-turn coding tasks, every model scored worse on these against-prior rules, i.e. once instructions pushed against its natural defaults.
在12个前沿模型和60个多轮编码任务中,每个模型在这些违背先验的规则上得分都更低,即一旦指令推动其偏离自然默认行为。
– arxiv. org/abs/2608.11727
– arxiv.org/abs/2608.11727
Title: "Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents"
标题:“Harness-IF:评估编码代理中跨指令表面的指令遵循能力”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力