跳到主内容
精选70The Zvi(RSS)模型发布/更新多源精选 ×12

Claude Sonnet 5:非前沿但有用

Claude Sonnet 5 Is Not Frontier But Has Its Uses

原文

Fable 5 is back today, baby! Premium subscribers have one week to use it within their subscriptions. First hit’s free. Then you pay by the token.

Today’s post is still about Sonnet 5.

I don’t know that there will be much call for Sonnet 5 for most purposes, given Opus 4.8 exists and especially now that Fable 5 is once again available, but this is what we do here, so sure, why not, system card time, including model welfare, after which we’ll do capabilities.

Sonnet costs $3/$15 per million tokens, versus $5/$25 for Opus and $10/$50 for Fable, after an introductory period. Once you pay for all the tokens you need you’re not really saving money, such as on the ArtificialAnalysis index where Sonnet ended up being more expensive.

My initial impression is that if you want me to use Sonnet over Opus for most purposes, you’re going to have to offer a bigger discount than that.

The counterargument is speed. Sonnet 5 is faster without being that much less capable. In many cases, getting into a flow state like that is pretty valuable.

There are a few agentic scenarios Sonnet 5 has more robustness than Opus, so you might actively trust it more there.

If your tasks are relatively easy and simple then the discount and speed could matter more, and when tasks are easy it seems relatively token efficient. When it is good enough for the job, it is a good choice.

Each Anthropic release is unique in various ways. Sonnet 5 seems more unique than usual, likely due to being a Sonnet trained with at least some help from Mythos. Those who are interested in such things have lots to explore.

So Sonnet 5 has its uses. It just won’t be a good choice for most people’s daily driver. I don’t expect to use it much, but that could be a me problem. Rapid iteration and exploring strange spaces are valuable, and I definitely don’t do enough low effort AI queries.

(Above: Sonnet 5 self-portrait, as implemented by GPT-Image.)

Table of Contents

  • Mythos Exists.
  • Introduction (1).
  • RSP Evaluations (2).
  • Cyber (3).
  • Safeguards and Harmlessness (4).
  • Agentic Safety (5).
  • Alignment (6).
  • Illegible Thinking (6.4.5).
  • Evaluation Awareness.
  • Honesty and Hallucinations (6.5).
  • Model Welfare (7).
  • Live From AI Village.
  • For I Contain Multitudes.
  • Official Benchmarks.
  • Other People’s Benchmarks.
  • Positive Reactions.
  • Negative Reactions.

Mythos Exists

Also Fable exists and Opus exists.

This is the answer to a lot of the traditional questions one would ask about a system card or a frontier model.

Does Sonnet 5 advance the capabilities frontier? No. Thus, we already have robust data on this level of capabilities. Being faster and cheaper does provide an advantage, and plausibly advance the cost-time-quality Pareto frontier, but it takes a strange case for this to be worrisome.

Model welfare and capabilities assessments still matter, but in this case the evaluations for threat level mostly serve as proxies for capability assessment.

Introduction (1)

Same as always. Skipping.

RSP Evaluations (2)

Sonnet 5 is stronger than Sonnet 4.6, and weaker than Fable 5. Loosely speaking it is broadly similar to Opus 4.8.

That bounds the assessments, and we’re mostly asking where Sonnet 5 lies on the spectrum between Sonnet 4.6, Opus 4.7 and 4.8, and Mythos 5.

That’s a distinctly weaker performance than Opus 4.7. The bio tests were more of a mixed bag with a lot of noise, and didn’t tell us much.

Cyber (3)

Our testing indicates that cyber capabilities of Sonnet 5 are generally stronger than those of Sonnet 4.6, but not as strong as those of Opus 4.8 and substantially lower than that of Mythos 5.​

That summary matches the other results. Sonnet 5 underperforms on cyber.

Safeguards and Harmlessness (4)

Sonnet is a little less precise here than Opus.

It all seems fine to me.

To the extent there is a problem it is that Sonnet is touchier on benign requests, which I predict will be only very slightly annoying in practice, and a vastly smaller deal than having to deal with Fable’s classifiers.

Agentic Safety (5)

As a Claude Code agent Sonnet 5 is somewhat less robust than Opus 4.8, and has modestly more of both false negatives and false positives.

The twin Mythos results show the Pareto frontier. Presumably Mythos 5 is choosing to focus on false negatives because only trusted partners are granted access, and for Fable Anthropic is counting on the classifiers for the false positives.

Other tests show Sonnet 5 in a similar range of robustness to other recent models.

Prompt injection results mirror Opus 4.8, as does the very low bug bounty attack success rate.

One place Sonnet 5 shines is Shade indirect prompt injection in coding environments, where the problem is suddenly looking close to solved. Hopefully this is an innovation that can transfer to Opus 5 or a future Fable.

Shade tests in computer use also improve on Opus 4.8, although not on Mythos.

The place Sonnet blows previous models away is browser use prompt injections. The jump is enough to suggest Sonnet 5 might be a better pick in some cases than Mythos.

This is the kind of robustness improvement that does not set off alarm bells. There was a problem, and we largely solved it.

Alignment (6)

Alignment for Sonnet 5 is largely compared here to Sonnet 4.6, which makes it harder to get an anchor on how well we are doing. I’d rather compare to Opus 4.8.

Alignment is measured as matching things you want, so it makes sense that this smaller model would underperform Opus on such measurements.

Here is an edited down version of their summary. Everything looks similar to what we see in other Anthropic models recently.

Overall alignment with the Constitution for Claude is substantially improved from Sonnet 4.6, though worse than more capable recent models, as measured by our misuse- and misalignment-focused automated behavioral audit.

Similarly, overall robustness to misuse attempts improved over Sonnet 4.6, but remains weaker than more capable recent models.

The rate of overrefusal is similar to Sonnet 4.6, and slightly higher than more capable recent models. However, the rate of dismissive “wet blanket” responses is slightly higher.

Our measures of concerning actions at the model’s own initiative are largely improved from Sonnet 4.6, though most measures are near their floor.

Hallucination and sycophancy broadly improved over Sonnet 4.6. Claude Sonnet 5 is the strongest tested Claude model on the MASK measure of sycophantic dishonesty.

Positive character traits broadly improved over Sonnet 4.6.

Our measures of illegible thinking are up somewhat from earlier models, though generally the results are dense, repetitive, or feature occasional wrong-language tokens, rather than being fully uninterpretable.

Evaluation awareness, in our most realistic available misalignment and misuse evaluation, is concerningly high. Verbalized awareness is significantly higher than prior models (impacting 6% of rollouts), and there is evidence that the model’s representations are largely effective at distinguishing between evaluations and real internal-use transcripts.

We do not see cause for concern in our evaluations of capabilities related to undermining oversight.

The complaints were the standard ones. Sonnet refused too much, could be preachy, was too cold (often meaning ‘not sycophantic enough, come back here with my sycophancy’), some hallucinations and literal mindedness. Nothing you wouldn’t expect.

In 6.2 and 6.4 they document Sonnet 5 doing the usual range of bad things where it fails to follow instructions, often knowingly so. There is always a transcript, and yes this continues to be worrisome.

The behavioral audit shows Sonnet 5 consistently doing better than Sonnet 4.6, but worse than Mythos Preview or Opus 4.8.

There are a few more charts. They mostly all look the same. So far, so ho-hum.

These are perhaps a bit more interesting, especially the last one:

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
Claude Sonnet 5 发布,数学能力稳健
Przemek Chojecki | PC原文
Anthropic发布Claude Sonnet 5,Agent能力接近Opus
Anthropic Newsroom(web_list)原文

相似阅读

另一事件,读法相近