跳到主内容
@wquguru
精选88METR(RSS)模型发布/更新多源精选 ×12

METR评估Claude Opus 5.5:AI研发加速有限,非范式跃迁

Summary of METR's predeployment evaluation of Claude Opus 5.5

原文
发到 X
推荐理由

第三方权威评估确认Opus 5.5为渐进升级而非范式跃迁,直接验证了当前大模型在自动化科研上的真实天花板,建议关注其技术边界。

Note on independence: This evaluation was conducted under an unpaid agreement for AI R&D assessment.1 We drafted the initial summary, and then Anthropic had the opportunity to review and edit the text. We signed off on this final text from the Claude Opus 5.5 system card.

独立性说明:本评估是在一项用于 AI 研发评估的无薪协议下进行的。1 我们起草了初步摘要,随后 Anthropic 有机会对文本进行审阅和编辑。我们签署了 Claude Opus 5.5 系统卡片的最终文本。

Our preliminary evaluation focused on how Claude Opus 5.5 might impact AI R&D, mainly based on its capabilities on difficult, long-horizon tasks. The main claims we attempt to assess in this report are: (A) would AI R&D at Anthropic now be dramatically accelerated by using Claude Opus 5.5; and (B) was AI R&D at Anthropic already dramatically accelerated due to AI during the development of Claude Opus 5.5.

我们的初步评估侧重于 Claude Opus 5.5 可能对 AI 研发产生的影响,主要基于其在困难、长周期任务上的能力。本报告试图评估的主要主张是:(A) 使用 Claude Opus 5.5 是否会大幅加速 Anthropic 的 AI 研发;以及 (B) 在 Claude Opus 5.5 开发期间,Anthropic 的 AI 研发是否已因 AI 而大幅加速。

Note that our work was oriented around collecting evidence related to AI R&D capabilities but was not meant to verify claims about compliance with any specific threshold from Anthropic’s policies. This report summary also does not attempt to assess whether Claude Opus 5.5 has or does not have particular alignment properties.

请注意,我们的工作围绕收集与 AI 研发能力相关的证据展开,但并非旨在验证关于符合 Anthropic 政策特定阈值的声明。本报告摘要也并未尝试评估 Claude Opus 5.5 是否具有或不具有特定的对齐属性。

Summary of evidence

证据摘要

We conducted a preliminary evaluation of Claude Opus 5.5 informed by:

我们对 Claude Opus 5.5 进行了初步评估,依据包括:

  • Capability testing, conducted via API access granted over a period of 10 business days. We used five tasks for this testing:
  • Budget NanoGPT Speedrun, a constrained version of the popular NanoGPT Speedrun competition for AI R&D.
  • Language Model Conceptual Argumentation (LMCA), a conceptual reasoning dataset described in A dataset of rated conceptual arguments (Cooper et al., 2026).
  • Train a Program, a task measuring an agent’s ability to train a machine learning model that replicates the behavior of a piece of software;
  • Gaming Bot, a task where an agent has to write a Python program capable of playing a video game, given a non-visual API to control the game.
  • Sunlight, a task measuring an agent’s ability to conduct open-ended research and write a corresponding report.
  • Background information about trends and capabilities of previous models, especially as reported in our recent Frontier Risk Report.
  • Additional information from Anthropic about Claude Opus 5.5’s capabilities, including responses to a questionnaire inquiring about model capabilities and control factors, and an interview with an Anthropic researcher.
  • A highly experimental and preliminary report from a separate METR assessment of AI R&D acceleration inside Anthropic. This report was written by a separate METR team with elevated access compared to the team that wrote this content. As a result of this difference in access, the separate METR team shared its conclusions with us, but was not able to share the supporting evidence or details of their reasoning.
  • Thus we use the evidence in the provided AI R&D report as an input to our assessment but do not argue directly in defense of its claims.
  • We consider its inclusion here a trial version of a more holistic assessment process. We expect further public outputs from this separate investigation in the coming weeks.
  • 能力测试,通过为期 10 个营业日的 API 访问权限进行。我们使用了五项任务进行测试:
  • 预算 NanoGPT Speedrun,这是针对 AI 研发的流行 NanoGPT Speedrun 竞赛的一个受限版本。
  • 语言模型概念论证(LMCA),这是一个概念推理数据集,详见《评级概念论证数据集》(Cooper 等人,2026)。
  • 训练程序,这是一项衡量智能体训练机器学习模型以复制软件行为能力的任务;
  • 游戏机器人,这是一项要求智能体编写 Python 程序以玩视频游戏的任务,给定一个非视觉 API 来控制游戏。
  • Sunlight,这是一项衡量智能体进行开放式研究并撰写相应报告能力的任务。
  • 关于先前模型趋势和能力的背景信息,特别是我们在最近的《前沿风险报告》中报道的内容。
  • 来自 Anthropic 关于 Claude Opus 5.5 能力的额外信息,包括对询问模型能力和控制因素的问卷的回复,以及与 Anthropic 研究人员的访谈。
  • 来自 Anthropic 内部 AI R&D 加速情况的另一项 METR 评估的高度实验性和初步报告。本报告由一个与撰写此内容的团队不同的 METR 团队编写,该团队拥有比原团队更高的访问权限。由于这种访问权限的差异,独立的 METR 团队向我们分享了其结论,但未能分享支持性证据或其推理细节。
  • 因此,我们将提供的 AI R&D 报告中的证据作为我们评估的输入,但并不直接为其主张进行辩护。
  • 我们认为将其纳入此处是一个更全面评估流程的试验版本。我们预计这项独立调查将在未来几周产生更多的公开成果。

Conclusions

结论

We divide our conclusions into two categories: those that relate to (A) the level of acceleration of AI R&D of which Claude Opus 5.5 is capable, and (B) the level of acceleration of AI R&D due to AI that went into the development of Claude Opus 5.5. Based on the available evidence, we arrive at the following conclusions:

我们将结论分为两类:(A) 与 Claude Opus 5.5 能够实现的人工智能研发加速水平相关的结论,以及 (B) 在 Claude Opus 5.5 开发过程中因人工智能而带来的 AI R&D 加速水平的结论。基于现有证据,我们得出以下结论:

(A) We believe that acceleration from this model would be slightly higher than for Fable 5.1, but that this model is unlikely to be able to fully automate AI R&D. This is because:

(A) 我们认为该模型带来的加速效应将略高于 Fable 5.1,但该模型不太可能完全自动化 AI R&D。这是因为:

  • We are reasonably confident that Claude Opus 5.5 does not represent a huge leap in AI R&D capability above Fable 5.1, but it likely represents a modest improvement upon Fable 5.1.
  • Claude Opus 5.5 improves upon Fable 5.1 across both verifiable tasks (Budget NanoGPT, Gaming Bot) and harder-to-verify tasks (LMCA, Sunlight), but:
  • In the questionnaire and researcher interview, Anthropic claims that Claude Opus 5.5 continues the Mythos-level trend on Anthropic ECI.
  • Claude Opus 5.5 is an incremental improvement above Fable 5.1 on our quantitative evaluations, rather than a discontinuous jump.
  • Claude Opus 5.5 still has qualitative weaknesses that an expert human is unlikely to exhibit when solving hard, long-horizon tasks or doing open-ended reasoning.
  • This is highly uncertain, but we expect that full automation of AI R&D will require large improvements in foresight, prediction, creating one’s own feedback loops, and generally other skills that might typically be referred to as researcher “judgement” or “taste”.
  • The evidence we have does not suggest that Claude Opus 5.5 represents a large improvement over Fable 5.1 in these “judgement” skills.
  • At the same time, we believe that Claude Opus 5.5 is still likely to noticeably accelerate researchers and automate limited aspects of R&D. For instance, we expect that Claude Opus 5.5 likely provides slightly higher productivity uplift than Fable 5.1.
  • Furthermore, frequent, incremental improvements on AI R&D ability are still consistent with a rapid overall rate of progress on AI R&D ability, but the data we have is insufficient for distinguishing consistent, accelerating, or decelerating rates of improvement. It is unclear how many continued incremental quantitative improvements of this size need to be made before qualitative changes in the AI R&D process can occur.
  • 我们有合理的信心认为,Claude Opus 5.5 并未代表相对于 Fable 5.1 在 AI R&D 能力上的巨大飞跃,但它可能代表了相对于 Fable 5.1 的适度改进。
  • Claude Opus 5.5 在可验证任务(Budget NanoGPT、Gaming Bot)和更难验证的任务(LMCA、Sunlight)方面均优于 Fable 5.1,但是:
  • 在问卷和研究者访谈中,Anthropic 声称 Claude Opus 5.5 延续了 Anthropic ECI 上的 Mythos 级别趋势。
  • 在我们的定量评估中,Claude Opus 5.5 是相对于 Fable 5.1 的渐进式改进,而非不连续的跳跃。
  • Claude Opus 5.5 仍然存在定性弱点,专家人类在解决困难、长周期任务或进行开放式推理时不太可能出现这些弱点。
  • 这一点具有高度不确定性,但我们预计,实现 AI R&D 的全面自动化需要在预见力、预测能力、创建自身反馈回路以及通常被称为研究者“判断”或“品味”的其他技能方面取得重大改进。
  • 我们所掌握的证据并不表明 Claude Opus 5.5 在这些“判断”技能方面相较于 Fable 5.1 有大幅改进。
  • 同时,我们相信 Claude Opus 5.5 仍可能显著加速研究人员的工作,并自动化研发(R&D)的某些有限方面。例如,我们预计 Claude Opus 5.5 带来的生产力提升幅度可能略高于 Fable 5.1。
  • 此外,AI 研发能力的频繁、渐进式改进仍然与 AI 研发能力整体快速进步的趋势一致,但我们现有的数据不足以区分这种改进是持续稳定、加速还是减速。目前尚不清楚需要进行多少次此类规模的渐进式定量改进,才能引发 AI 研发过程的质性变化。

We also made use of an additional source of information which we are not able to disclose at this time. We used this source to understand AI R&D at Anthropic, but it did not relate to METR’s understanding of this particular model.

我们还利用了另一信息来源,但目前无法披露该信息。我们使用该来源来了解 Anthropic 的 AI 研发情况,但它与 METR 对该特定模型的理解无关。

(B) We believe that the development of this model was at least somewhat accelerated by AI but is unlikely to have been dramatically accelerated by AI. This is because:

(B) 我们认为该模型的开发至少在一定程度上因 AI 而加速,但不太可能因 AI 而大幅加速。原因如下:

  • The estimate provided by the preliminary AI R&D report is “~1.5X overall acceleration in capabilities due to AI (i.e. 1.5 years in 1 year), with perhaps 30% chance of 2X acceleration.”
  • Note that because the preliminary report did not specify the time period for this estimate, it is unclear whether this estimate applies to the development of Claude Opus 5.5 or another period.
  • Our understanding of Claude Opus 5.5’s capabilities suggest it is an on-trend improvement over Fable 5.1.
  • This is based on our internal capability evaluations suggesting an incremental improvement over Fable 5.1, as well as information provided by Anthropic in the questionnaire and researcher interview, which indicates that Claude Opus 5.5 continues the AECI trend from Mythos Preview onward.
  • Our AI R&D assessment agreement with Anthropic includes a transparency provision that lets us disclose certain facts publicly without Anthropic’s consent: the existence and general terms of the review and redaction process, whether Anthropic exercised its redaction rights, whether any of our findings depended significantly on redacted information, the level of access we were given, and the confidentiality framework that applied to the evaluation. ↩
  • 初步 AI 研发报告提供的估计值为“由于 AI 的作用,整体能力提升加速约 1.5 倍(即 1 年内完成原本需 1.5 年的工作),其中加速 2 倍的概率约为 30%。”
  • 请注意,由于初步报告未明确该估计值所涵盖的时间段,因此尚不清楚该估计值是否适用于 Claude Opus 5.5 的开发,还是指其他时期。
  • 我们对 Claude Opus 5.5 能力的理解表明,它是相对于 Fable 5.1 的符合趋势的改进。
  • 这是基于我们内部的能力评估,显示其较 Fable 5.1 有渐进式提升;同时也基于 Anthropic 在问卷和研究人员访谈中提供的信息,这些信息表明 Claude Opus 5.5 延续了自 Mythos Preview 以来的 AECI 趋势。
  • 我们与 Anthropic 签订的 AI 研发评估协议中包含一项透明度条款,允许我们在未经 Anthropic 同意的情况下公开某些事实:审查和删减流程的存在及一般条款、Anthropic 是否行使了删减权、我们的任何发现是否严重依赖于被删减的信息、我们获得的访问权限级别,以及适用于本次评估的保密框架。↩

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件