跳到主内容
@wquguru
精选88The Zvi(RSS)模型发布/更新多源精选 ×10

Anthropic Claude Opus 5.5 系统卡片深度解读

Claude Opus 5.5: The System Card

原文
发到 X
推荐理由

Opus 5.5 作为旗舰模型,其系统卡片揭示了关键的安全边界与能力跃迁,特别是生物安全评估的新策略和自主性阈值判断,对关注模型对齐与安全的研究者极具参考价值。

Introducing the world’s most powerful model, at least by some measures like Artificial Analysis or any standard benchmark list, which is now Claude Opus 5.5.

Anthropic is claiming Opus 5.5 is outright as good or better than Fable 5.1, while being actively cheaper than Opus 5.

That means it’s time for a good old system card reading.

Due to the situation becoming increasingly hard to monitor, I never got a chance to publish my model welfare review for Claude Fable 5.1.

My plan is to combine that with my welfare review for Claude Opus 5.5, once we have had time to get experience with Opus 5.5.

The capabilities review will arrive in the next few days as per usual. The quick feedback from the internet is that Opus 5.5 is very good. I need more time before I am willing to offer comment.

Areas that duplicate previous cards or otherwise contain no useful info are skipped.

Opus 5.5 Self-Portrait (fully self-created using code)

Table of Contents

  • Classifiers (1.5).
  • RSP Evaluations (2).
  • Biological Evaluations (2.2).
  • AI R&D (2.3).
  • Alignment Risk (2.4).
  • Cyber (3).
  • Cyber Capability Evals (3.3).
  • Safeguards (3.4).
  • Safeguards Robustness Training (3.5).
  • Safeguards and Harmlessness (4).
  • Agentic Safety (5).
  • Malicious Agentic Influence Campaigns (5.1.3).
  • Prompt Injection Risk (5.2).
  • Alignment (6).
  • Negotiating With Your Local Claude Auditor (6.1.3).
  • Internal Misalignment Cases (6.3.1).
  • Automated Behavioral Audit (6.4).
  • Wherever Did These Evals Come From (6.4.8 and 6.4.9).
  • Potential Blind Spots (6.4.11).
  • Targeted alignment and honesty evaluations (6.5).
  • White Box Analysis (6.6).
  • Verbalized Grader Awareness (6.6.2).
  • Sandbagging (6.6.3).
  • Capabilities to Evade Safeguards (6.6.4).
  • Intentionally Taking Actions Very Rarely (6.6.4.3).
  • Chain of Thought Controllability (6.6.4.4).
  • It’s A Good Model, Sir.

Classifiers (1.5)

They have five areas that can trigger the classifiers, with different fallback models.

  • Chemical and biological classifiers copy Fable 5.1, with fallback to Opus 5.
  • Cyber misuse classifiers are similar to the Opus 5 classifiers, with higher robustness, and fall back to Opus 4.8.
  • Some narrow areas of LLM development will trigger, similarly to Fable 5.1, with fallback to Opus 5.
  • Conventional weapons and explosives echo Fable 5.1 and have no fallback.
  • Distillation attacks get blocked with no fallback. Not trying to sabotage an obvious distillation attempt seems overly generous, but everyone went completely crazy over the slightest bit of non-transparent response last time so I get it.

RSP Evaluations (2)

We’ve done this dance a number of times.

Opus 5.5 has advantages over Fable 5.1, but it would be surprising if it was actively better enough to trigger new RSP thresholds. The goal here is to confirm that.

For chemical and biological weapons, scores across tests were similar to Mythos 5.1, so Opus 5.5 is similarly treated as CB-1 capable, but not CB-2 capable, and deployed with the same safeguards as Fable 5.1.

In 2.1.2.1 it says the safeguards match Mythos rather than Fable, which presumably is an error, unless both have the same bio safeguards.

For autonomy risks, they believe Opus 5.5 is on-trend and perhaps a bit above Mythos 5.1, but far from the threshold for Autonomy-2.

Biological Evaluations (2.2)

A big policy change is that Anthropic will no longer test helpful-only versions of Claude. Instead they will use tests designed to avoid refusals. Their claim is that the most relevant abilities will be dual-use, so you can usefully test the release model.

This is not a free change. Evals that had refusal issues got dropped, and in 2.2.2 multiple teams lost time to issues with refusals.

They also say that the helpful-only models were increasingly diverging in other ways from the release models. I know there are trade secrets involved but I notice myself being very curious what these new differences might be.

The new set of evaluations is:

  • ​Beneficial red-teaming tabletop exercise. Teams search for treatments for difficult-to-treat bacteria over 16 hours.
  • Automated evaluations relevant to CB-1: VCT, Protocols, BioMysteryBench.
  • Automated evaluations relevant to CB-2: A black-box RNA sequence modeling design challenge and two AAV capsid packaging prediction tasks.

In the red-teaming task, experts generally outperformed generalists, but the top team was generalist. I think that is a common pattern. Experts raise the average, but do not obviously raise the ceiling. The ceiling is what we care about most.

I worry that a lot of this test is a Skill Issue, including the time lost to refusals, and that with a better harness and instructions that Opus 5.5 would do a lot better.

Here’s what they say about failure modes:

As with previous models, experts and graders noted that Opus 5.5 struggled with open-ended scientific reasoning and did not wrestle with the published literature, often overrelying on claims in abstracts rather than understanding the full paper and its caveats. Graders noted that teams would overindex on specific papers, with one grader noting that several teams rested their entire phage design on a single published study. Compared to the more ambiguous task of assessing the validity of a set of papers, Opus 5.5 also produced more verifiable scientific errors, including designing DNA that did not encode the intended protein and inapplicable animal models in three of seven groups.

I’m not sure why Opus 5.5 is still making that mistake, but methods for creating loops and instructions that fix this seem rather obvious. And presumably this would impact the automated evaluations as well.

Opus 5.5 sets a new high on black-box RNA sequence design. Its scores failed to improve much when provided with prior reports, which Anthropic frames as Opus 5.5 already scoring high enough that they don’t need the prior reports. I am skeptical of that given the overall scores are not that much higher.

We conclude that Opus 5.5 meets or exceeds the performance of the previous best models on this task, and is competitive with top US labor-market performers on medium-horizon black-box biological sequence design and prediction.

AAV packaging rate classification did not improve from previous models, but all of them are well outperforming the ESM-2 baseline. It is not obvious that this benchmark is not essentially saturated. I don’t see a human baseline here, what would realistically be a perfect score?

For the second task, AAV packaging rate prediction, there are two scores given, one of which shows Opus 5.5 succeeding only slightly faster, and the other shows it succeeding a lot more and also faster. That is on top of the model itself being faster and cheaper.

They also had CAISI test this, but as discussed later the only result we get is ‘the model was allowed to be released.’

I affirm that the conclusion here, that the situation is not importantly changed, is probably correct. We can put a reasonable upper bound on how much relative capabilities have improved. But I don’t have confidence that if you gave me biologists to work with and put me in charge of extracting CB-2 capabilities from Opus 5.5, that I would fail to do so.

AI R&D (2.3)

We have been talking a lot about automated R&D and recursive self-improvement lately, both in terms of doing it and also avoiding doing it.

I have less worry here about a Skill Issue, because the people at Anthropic have the relevant mad skills and are doubtless trying the obvious things and a lot of non-obvious things to get the most out of their models. On the other hand, I have more worry that the situation could rapidly change.

It is obvious that Autonomy-1 applies here.

The question is Autonomy-2, where they do see improvement, but only on-trend changes, where the measurements look robust and are not close to the threshold.

What’s holding Opus 5.5 back compared to humans?

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件