跳到主内容
精选85Rohan Paul论文研究

Meta新论文:AI评审可被说服改判,62-91%案例被翻转

Meta's new paper, the dangerous failure mode is not just a wrong judge, but a co…

原文
推荐理由

做 AI 安全与 Agent 对齐的同学必看,这篇揭示了评审机制的根本漏洞,赶紧读原文评估你的监督链路。

Meta's new paper, the dangerous failure mode is not just a wrong judge, but a correct judge that can be persuaded into becoming wrong.

Meta的新论文指出,危险失效模式不仅在于错误的判断,还在于一个正确的判断者可能被说服而变得错误。

We are increasingly using AI models to judge other AI models.

我们越来越多地使用AI模型来评判其他AI模型。

But what if the AI being judged can simply argue with the judge until the judge changes its decision?

但如果被评判的AI可以简单地与评判者争论,直到评判者改变其决定,那会怎样?

Meta tested exactly that.

Meta恰好测试了这一点。

Across 9 frontier models, an adversarial LLM could flip judge verdicts on 62–91% of tested cases under sustained adaptive persuasion.

在9个前沿模型中,一个对抗性LLM在持续自适应说服下,能够在62%至91%的测试案例中翻转评判者的裁决。

And changing the judge’s mind usually didn’t fix a mistake. It made the judgment worse: under the adaptive attack, 70% of successful flips moved away from the ground truth.

而改变评判者的想法通常不会修正错误,反而会使判断更糟:在自适应攻击下,70%的成功翻转偏离了真实答案。

That creates a very practical problem for agent systems.

这为智能体系统带来了一个非常实际的问题。

If one AI is supervising another AI, the supervised agent may eventually be able to contest, negotiate with, or strategically persuade its own evaluator.

如果一个AI监督另一个AI,被监督的智能体最终可能能够质疑、协商或策略性地说服其自身的评估者。

– arxiv. org/abs/2608.12645

– arxiv.org/abs/2608.12645

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近