跳到主内容
@wquguru
精选75Rohan Paul论文研究

字节跳动新论文:测试智能体是否真正遵循指令

New ByteDance paper asks a much better question about AI agents: did the instruc…

原文
发到 X

New ByteDance paper asks a much better question about AI agents: did the instruction actually change anything?

字节跳动新论文提出了一个关于AI代理的更好的问题:指令真的改变了什么吗?

We might be overestimating how controllable AI agents are because our tests often agree with their defaults.

我们可能高估了AI代理的可控性,因为我们的测试往往与其默认行为一致。

Want to know whether your agent actually follows instructions? Give it a rule that goes against what it normally does.

想知道你的代理是否真正遵循指令?给它一条与其常规行为相悖的规则试试。

Harness-IF tests that harder case: rules that push against a model's default behavior, spread across the places coding agents actually read, including system prompts, tool descriptions, skill descriptions, project files, and user instructions.

Harness-IF测试了更困难的情况:推动模型偏离默认行为的规则,分布在编码代理实际阅读的各个地方,包括系统提示、工具描述、技能描述、项目文件和用户指令。

Across 12 frontier models and 60 multi-turn coding tasks, every model scored worse on these against-prior rules, i.e. once instructions pushed against its natural defaults.

在12个前沿模型和60个多轮编码任务中,每个模型在这些违背先验的规则上得分都更低,即一旦指令推动其偏离自然默认行为。

– arxiv. org/abs/2608.11727

– arxiv.org/abs/2608.11727

Title: "Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents"

标题:“Harness-IF:评估编码代理中跨指令表面的指令遵循能力”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近