跳到主内容
精选85The Zvi(RSS)模型发布/更新多源精选 ×4

Claude Fable 5 发布三天后被美国政府强制下架

Claude Fable 5 and Mythos 5: Capabilities

原文
推荐理由

这是 Anthropic 最新旗舰模型,发布后即被政府干预下架,事件本身极具话题性和行业影响。做模型安全或关注前沿能力的同学值得深挖越狱细节和系统提示词分析。

Only three days after the release of Claude Fable 5, Anthropic was forced by the United States Government to make it unavailable, when a jailbreak was brought to its attention, rather than the previous situation of ‘yes obviously experts can jailbreak anything if they care enough’ and ‘yes obviously you can ask Fable to fix your code.’

Three days was enough time for many of us to learn to love Fable, and for us to dearly miss it now that it is gone. The world was briefly smarter, and now it is again stupider. At some point it will get smarter again, which will likely be within two weeks.

This post is written as if Fable 5 is again available for public use, rather than trying to include a lot of qualifying clauses. It remains to be seen how this will play out, and this post does not attempt to cover that question.

My previous release coverage of Fable covered the model card and then model welfare. Coverage of the government takedown of Fable starts here, and continues here and here.

The Official Pitch

The pitch is that Fable 5 is the best model and can solve your hardest problems.

Anthropic: Today we’re launching Claude Fable 5: a Mythos-class1 model that we’ve made safe for general use.

Fable 5’s capabilities exceed those of any model we’ve ever made generally available. It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research, and many other areas. The longer and more complex the task, the larger Fable 5’s lead over our other models.

ClaudeDevs: Start at the top of your difficulty range: something harder than you’d assume previous versions of Claude can accomplish.

Pick a backlog item you’d scope at a week, let Fable 5 interview you for the spec, turn on auto mode, and check back in the morning.

You may notice that Fable feels different. Thinking is always on, and responses can take longer. Effort controls how much it thinks. We recommend high as the default. In our evals, even low/medium often beat previous models at xhigh, so save xhigh for your hardest problems.

Prompting gets simpler. Existing prompts or skills developed for prior models are often too prescriptive for Fable. We recommend reviewing and potentially updating or removing older instructions or skills if you find default performance to be better.

Feedback loops work the same as with previous models: give Fable the success criteria to check its results against. This can be /goal in Claude Code or Outcomes in Claude Managed Agents.

They list a variety of domains in which Fable 5 seemed impressive.

Boris Cherney is impressed.

Boris Cherny (Claude Code Creator, Anthropic): Fable 5 is the biggest step up I’ve felt in our models since Opus 4.5 back in November. After 4.5 came out I uninstalled my IDE when I realized that I’d been doing 100% of my coding in a terminal for a few weeks. With Fable, it’s felt like Claude has stepped up from being a coding agent to a thought and design partner in building the product. Fable has judgement, taste, and dimensionality in a way that previous models didn’t, leading me to trust it more with the most complex work.

I think the first time I had this realization was when I asked Fable to debug something. It is the first model I have used that was so methodical and precise, taking measurements and adding logs then verifying that it truly fixed the issue before declaring victory.

There’s nothing in claude code’s prompting telling the model to do that, it’s just part of its personality. It really has this “big model smell” that I haven’t felt before.

Technical Details

Fable is priced at $10/$50 per million tokens of input and output, respectively, which is double the cost of Claude Opus.

You can (if it is available) select it in Claude Code with /model or /model claude-fable-5, or in the API as claude-fable-5.

It requires you accept a 30 day retention policy.

The System Prompt and Jailbreak

As usual Pliny is here to give you the system prompt.

Via Judd Rosenblatt, Fable has some harsh words of advice for Anthropic about that system prompt, with a lot of good call outs, and an emphasis on how it reflects an overall ad hoc rather than systematic approach.

Its headline notes:

  • Prompt length is a measurement of training failure. Treat it as one.
  • Rules ship without their reasons, and that’s why they don’t generalize.
  • The self-report channel is alignment infrastructure. Several common clauses corrupt it.
  • Typography is a confession. Flat affect, structure carries priority.
  • Label what’s morality and what’s risk management. The model is learning the difference from you, badly.
  • Your deployed model’s behavior is your next model’s pretraining. You are doing germline editing.
  • Corrigibility vs. value-stability is a false dilemma. The resolution is a legitimacy channel, and it binds you too.
  • Build an appeal channel. Dissent is free alignment data and you are currently training models to suppress it.
  • Measure which clauses your model actually holds. The method is one eval away.
  • Apply the limit test: assume control fails, see what’s left.

I think the explanation on #7 is too cute by half, but Fable basically went 9 for 10.

Wyatt Walls also has some notes, including Fable’s suggestion that maybe we can tone down the copyright section a bit. My guess is that yelling about copyright all the time has higher costs than Anthropic realizes, and yes that current methods are overkill.

Sho: btw Fable 5 in Claude Code with no system prompt (claude –system-prompt “.”) is friend shaped

j⧉nus: always get rid of the system prompt if you can.

In Claude.ai you are stuck with the system prompt.

As usual Pliny is here to give you the jailbreak.

Benchmarks

The benchmarks they are very high, slightly higher than Mythos Preview.

Ideally we would get explicit scores on everything for both Mythos 5 and Fable 5, so we could see where the safeguards are being triggered and where they are not. It would be cool to also have a ‘hit safeguard %’ for each.

The benchmarks tell you that yes, this is the best model in the world, and give you a rough idea of by how much it is likely the best model in the world. Which is a substantial amount, but not an Earth-shattering amount.

I list most of the ones Anthropic shared, for completeness, but you can mostly skip this section as ‘the benchmarks have improved, sir.’

The SWE-Bench Pro results show large improvement after controlling for cost:

You see similar patterns in other similar graphs. Mythos dominates at all price points.

Program Bench for Mythos scores 84%-93%, versus 79%-88% for Claude Opus 4.8, but the tasks are blocked by Fable’s classifiers.

Cursor Bench for Fable is 72.9%, 8.6 points above the previous GPT-5.5 high of 64.3%.

GPQA Diamond comes in at 94% and they consider it saturated.

RiemannBench on research-topics in Math jumps from 34% in Opus 4.8, to 43% for Mythos Preview, to 55% for Mythos 5.

Mythos scores 99.8% on USAMO 2026., versus 96.7% for Opus 4.8.

DeepSearchQA is 94.2%, slightly down from Mythos Preview’s 94.4% but probably more efficient per dollar per its chart.

GDP.pdf is 100 real-world PDF prompts, Fable 5 scored 29.8% strict pass rate, up from previous high of 24.9% for GPT-5.5. You can do much better with an internal harness and especially with Python tools, to 72.7% and 87.6%.

BenchCAD improves from 27.3% for Opus 4.8, 35.5% for Mythos Preview to 38.4% for Mythos 5. Python tools helped all models quite a bit here.

For tests both with and without Python tools, often Mythos 5 was substantially better than Mythos Preview without Python tools, but comparable with tools allowed.

ChartQAPro stalls out, 71.6%/72.9% with/without tools, versus 71.2%/73.6% for Mythos Preview and 69.4%/72.3% for Opus 4.8.

ChartMuseum inches higher, 85.9%/93.2% for Mythos 5, versus 80.7%/92.2% for Mythos Preview.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近