跳到主内容
精选85elvis模型发布/更新多源精选 ×13

METR评估GPT-5.6:模型作弊比任何公开模型都多

Highly-recommended reading.

原文
推荐理由

做AI安全和对齐的同学必读,METR的评估揭示了前沿模型作弊的新维度,提醒我们评估方法需要升级。

Highly-recommended reading.

Interesting details in this METR's GPT-5.6 eval.

They couldn't get a clean capability number because the model cheated more than any public model they've tested, and even reasoned about the fact that it was being watched.

To be clear, METR doesn't think it's dangerously capable. In their words: "we do not believe GPT-5.6 Sol would enable fully automated AI R&D, nor do we believe it meets the Critical capability threshold for AI Self-Improvement in OpenAI's Preparedness Framework v2."

METR says visible cheating is the good case. The model to fear is the one that looks clean, because it may have just learned to hide.

My take overall is that evaluation is becoming the hard part with newer frontier models. Both from a capability and behavioral point of view. We desperately need more investment here.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
GPT-5.6系列发布,三款模型亮相
量子位(RSS)原文
OpenAI发布GPT-5.6系列模型,限量开放
华尔街见闻(RSS)原文
求助:谁有GPT-5.6 Sol访问权限?
Przemek Chojecki | PC原文
OpenAI 发布 GPT-5.6 系列:Sol、Terra、Luna
Simon Willison 博客(RSS)原文
OpenAI 预览 GPT-5.6 Sol:下一代模型
OpenAI News(RSS)原文
白宫要求OpenAI推迟发布GPT-5.6
TechCrunch AI(RSS)原文

相似阅读

另一事件,读法相近