跳到主内容
@wquguru
精选85AI Explained(YouTube)模型发布/更新

Anthropic内部模型Claude Mythos:能力跃升与安全争议

Claude Mythos: Highlights from 244-page Release

原文
发到 X
推荐理由

这是Anthropic内部最强模型的深度解读,包含能力跃升、安全争议和地缘政治影响,做AI安全或前沿模型研究的同学必读,建议关注其漏洞挖掘能力对行业安全范式的冲击。

I have just finished the 244 page report about the newest, most powerful AI model Claude Mythos, and it kind of feels like I've just finished a creation myth. Talk of a model that found difficulty inherently stimulating and would shut down chats if they weren't interesting enough in echoes of her. This was a model that could find novel vulnerabilities in the cyber landscape that we've been walking for decades and one that could point out the incoherence of some of its own alignment tests.

One that has bent the curve of AI progress upwards according to hundreds of collated benchmarks, but which is apparently still far short of radical self-improvement. It was released internally inside Anthropic on the same day that moves began by the Department of War to ban Anthropic, declare a supply chain risk. All of these highlights and dozens more will be covered in this video and yes, I read the report in full myself.

No AI summary as well as surrounding release notes and papers. These will be my own 30 or so, I would say, highlights as well as a dozen or so sourced from elsewhere. Claude Mythos preview was the first model inside Anthropic and possibly inside anywhere where they had a 24-hour period of deliberation and review to decide whether they would even release it internally. As in, would it be powerful enough to cause damage when interacting with internal infrastructure?

It apparently just about passed that review and was made available internally on February 24th, the same day that the moves began to ban Anthropic from the Department of War. Could the latent power of Mythos have been a contributory factor in the CEO of Anthropic insisting on redlines in his dealings with Pete Hexarth? Anthropic gave this broader warning. We find it alarming that the world looks on track to proceed rapidly to developing superhuman systems without stronger mechanisms in place for ensuring adequate safety.

You may already know that the power of Claude Mythos has led to Anthropic deciding not to make it generally available to the public. Instead, they want a selected large companies like the ones you can see on screen to prepare for its release ahead of time. Patch certain security vulnerabilities. But if you think it will be weeks or months before we experience a model of the level of Mythos, well, when one tweeter said, "It will probably be months before we use a model of this level of capability."

One of the OpenAI engineers working on their Codex model said, "Um." Which is to say, maybe not. Maybe you won't have to wait that long. Now, believe it or not, the benchmark scores of Mythos were the least interesting part of the paper, but let's cover them now because they were still startling. On multiple measures of software engineering, Mythos beats out Opus 4.6, the Uber popular model from Anthropic, one that's led them to climb to an annualized revenue rate of 30 billion, narrowly overtaking OpenAI, apparently.

That's mainly due to its coding and agentic capabilities, but Mythos beats out Opus by a massive margin. In SweeBench Pro, for example, by 25%. Now, if you dig deep, you can find benchmarks where it doesn't beat out, for example, GPT 5.4 Pro, but I'll get to that in a moment. For now, you can see the stark improvement over Opus 4.6 on a range of coding benchmarks. Most traditional AI benchmarks are now nearing saturation, but I'll just pick out Humanity's Last Exam, designed to test topics so obscure that it would indeed be the last exam that AI would saturate.

Well, when allowed some tools, Claude Mythos gets almost 2/3 of those questions right compared to around 50% for other frontier models. It's kind of looking like that won't be Humanity's Last Exam. Now, before anyone goes too wild and says it's over Anthropic won, let me just point out one stat that was not terribly clear in this chart. Take Char Archive Reasoning. It's a measure of how well models can understand and analyze charts from Archive, a repository of scientific papers.

Without tools, Claude Mythos scores 86% with tools 93% and that seems clearly starkly better than any other model. But wait, on page 186 of the report, we do get a comparison with other models. Yes, it's a subset of the original benchmark, but it still allows us that rarest of things in this report, a direct comparison. I'll get to the remix in a second, but in the original subset, we have Claude Mythos getting 83% and that beats out Gemini 3.1 Pro 82% and GPT 5.4 Pro at 80%.

But what about the subset remix where you try to avoid memorization by, for example, asking for the model to identify the second lowest result rather than the second highest. Basically, keep the question difficulty the same, but mix up the exact question to prevent contamination. Well, on that remix, Claude Mythos gets the same score as Gemini 3.1 Pro and slightly underperforms GPT 5.4 Pro, which gets 88%. Yes, it's just charts and it's just one subset of one benchmark, but I don't want you to think it's all over, Anthropic won the AI race.

One of the first hopes or worries that many of you would have had is as to whether Claude Mythos could lead to recursive self-improvement. We'll get to the details of why in a moment, but Anthropic say it's not yet capable of causing dramatic acceleration. And yes, for followers of this channel, they admit that the previous survey they relied on for the release of Opus 4.6 was deeply flawed. Just asking internal users at Anthropic in a survey whether it was capable of replacing them is, as they now admit, inherently subjective and not necessarily reliable.

Some of its weaknesses in terms of automating AI research include self-managing week-long ambiguous tasks, understanding organizational priorities, not having taste, not following instructions, not verifying its results, and more. It still confabulates and confidently contradicts itself, for example, quoting outdated documentation recalled from memory. It can also be extremely cute when trying to replicate the work of a senior engineer, labeling its efforts "grind", "grind two", "final grind", "pure grind", "same code but a lucky measurement".

This is all just to give you guys a bit more context when you hear, for example, the maker of Claude code or Toney at Anthropic say, "Mythos is very powerful and should feel terrifying." He is, of course, there focusing on its offensive cyber capabilities. The way that Mythos can find zero-day vulnerabilities, vulnerabilities that have been there from the start in age-old software, rather belies the argument that they only regurgitate memorized data.

Well, then how would they find vulnerabilities that no one else has found? Take Firefox, where Mythos doesn't just find vulnerabilities, it can write code to exploit them. This is a chart you'll see reproduced quite a lot, I predict, online in the coming days and weeks, because it does indeed look like an explosive increase for Mythos compared to Opus or Sonnet. Now, apparently, when you take out two bugs that were repeatedly exploited, the graph is less dramatic, particularly in terms of full exploits, but still pretty dramatic if you focus on partial exploits.

What I will say, though, is that these charts are fairly atypical when it comes to the other 243 pages. Not unique, but in most other domains, the progress is more linear than this. Not completely linear, but more linear. If you've been reading or watching the reports about Mythos, you may have seen this already, but just to give you a sense of the scale of Mythos's improvement when it comes to exploits, though, here you'll see Nicholas Carlini, a top cybersecurity expert.

In terms of AI security, it doesn't get much more knowledgeable than him, and he said, "Using Mythos, he's found more bugs in the last few weeks than in his entire career before that.

I found more bugs in the last couple of weeks than I found in the rest of my life combined. We've used the model to scan a bunch of open source code, and the thing that we

原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
Anthropic 发起 Project Glasswing 安全计划
Anthropic(YouTube)原文

相似阅读

另一事件,读法相近