白宫叫停Claude Fable 5与Mythos 5部署
AI #173: AI Pauses
这是AI行业近期最重磅的政策事件:白宫直接动用出口管制叫停前沿模型部署,影响Anthropic旗舰产品,且涉及模型能力与安全监管的根本矛盾。从业者应密切关注后续进展,评估对自身业务和合规策略的潜在影响。
A lot of things are always happening. Only one story matters.
Claude Fable 5 and Claude Mythos 5 were shut down, by the White House, via an imposition of export controls at 5:23pm on Friday, wreaking all sorts of havoc.
There was then a scramble. Anthropic flew its people out to Washington, where they met with the Trump Administration on Monday, with hopes expressed that this could be quickly resolved.
What caused this? The Trump Administration said it was due to a jailbreak of Fable, which we now know they were told about by Amazon. They called Dario Amodei, who they complain did not take the issue sufficiently seriously. Rather than shutting down the model, he tried to explain why he saw no need to do that. This did not go well.
The ‘jailbreak’ turns out to be saying ‘fix this code,’ and the demo was getting Fable to find the same weaknesses that were easily identified by Opus 4.8 and GPT-5.5. As in, Fable is willing to work to fix security vulnerabilities if you give it a codebase. From this information and process, you could then figure out what the original bug in the code was, and exploit it, despite Fable refusing to to do that if you typed in ‘hack this server.’
The Trump administration now says that Fable can come back online when Anthropic ‘fixes’ this ‘jailbreak.’ That is of course impossible. This cannot be fixed. Your AI is either highly skilled at and capable of writing secure code, or it is not. You cannot draw this level of distinction between offensive and defensive capability.
The only ways to have this not allow you to route around the classifiers are either to have the classifiers not try to block similar requests in the first place, or to broadly take away Fable’s ability to code.
This is now day seven of this pause in the deployment of frontier AI capabilities.
We continue to be a little under even money for it to end by July 1.
Check the bold links above for my full coverage of that.
This post is mostly about everything else that is happening.
That includes some really cool things, such as MidJourney Medical announcing a new method of full body scanning with no health risks, no radiation and super high resolution, at very low marginal cost, that they hope to start deploying next year.
Last week Anthropic dropped some policy proposals. It seems quaint already, but I review those here.
Table of Contents
- Language Models Offer Mundane Utility. Ask AI about markets in everything.
- Language Models Don’t Offer Mundane Utility. Offer may be void in the EU.
- Huh, Upgrades. Usage limits get more generous.
- On Your Marks. We add AA v4.1, EvalEval and Opus Magnum.
- VirtueBench. We also get VirtueBench. Is your AI a good Augustine?
- Choose Your Fighter. Microsoft considering DeepSeek for Copilot.
- Papers, Please. Anthropic reserves the right to confirm your identity.
- Deepfaketown and Botpocalypse Soon. AI used by police to fabricate evidence.
- Goodhart’s Law Strikes Again. Have you considered minimizing costs?
- They Took Our Jobs. The situation is escalating quickly.
- The MidJourney Full Body Imaging Scanner. This is so cool.
- Introducing. GLM-5.2 talks big, Cursor trains a model, OpenRouter tricks.
- In Other AI News. Who gets how much value out of agentic coding?
- Show Me the Money. DeepSeek raises $7.5B at $50B.
- Bubble, Bubble, Toil and Trouble. Trying to steelman the bubble case.
- Quiet Speculations. Is customer optimization a threat to corporate profits?
- People Just Say Things.
- The Widened Path. DeepMind sees four ways to get to superintelligence.
- Scott Alexander Lays Out His AI Opinions. Now you know.
- Quickly, There’s No Time. Humans have been recursively self-improving.
- Policy On The AI Exponential. Dario writes another soft-pedaling essay.
- Anthropic Offers Two Policy Frameworks. An interesting timing choice.
- Obligations of Developers. These are not ambitious obligations, but sure.
- Societal Resilience Measures. Insufficient, but yes, obviously do these things.
- Economic Policy Framework. Gesturing towards redistribution.
- White House Pauses AI Deployment. This is our new reality.
- The Once And Future Fable. An attempt to build a sane formal process.
- How To Fix This Code. Can’t have a jailbreak if no one is in jail.
- The End of Privacy. Export controls as a path to broad identify verification.
- AIs Have Preferences. What tier are you in?
- The Quest for Sane Regulations. Congress moves to limit abuse of process.
- Chip City. NAACP is latest to attack data centers.
- The Week in Audio. Nate Soares on Will Cain, Dario Amodei on Bloomberg.
- Rhetorical Innovation. Have you considered the ‘wrong hands’ could be digital?
- Aligning a Smarter Than Human Intelligence is Difficult. Cheat, cheat, cheat.
- People Are Worried About AI Killing Everyone. The AIs.
- The Lighter Side. The news never stops.
Language Models Offer Mundane Utility
Ask your AI how to ask your AI.
Build markets in everything, in this case in hay.
Language Models Don’t Offer Mundane Utility
KPMG report on benefits of AI contained AI hallucinations.
Siri AI will not be coming to Europe due to the Digital Markets Act, since if it ships then every rival agents must get the same access to data as Siri. Apple is unwilling to offer that, for obvious security reasons.
Huh, Upgrades
Codex adds ability to bank its limit resets, which is a lot like saying you get credits over time that don’t expire, with different labels. It also is a de facto price drop and very customer friendly, so I approve.
Anthropic indefinitely rolls back disallowing programmatic use of its Claude Code subscription quotas. In a sufficiently long run this is not a sustainable cost structure, but for now it seems good.
On Your Marks
EvalEval Coalition will assemble all the evals in one place and tells you how each was made and how much you can trust them. When I checked the actual results were not ready yet.
Opus Magnum, a game high on my wish list, becomes a new benchmark.
Rob Haisfield: Are AI agents shape rotators? In this new benchmark, we let the models play campaign puzzles in Opus Magnum, a puzzle game by @zachtronics .
Ironically, Claude Opus 4.8 performed poorly, being beaten by GPT-5.5, Gemini 3.5 Flash, and GLM 5.2. Claude Fable 5 crushed them all.
No language model solved all 36 puzzles. Fable 5 and GPT-5.5 performed best, with GLM 5.2 as the best open weights model. No model beat a human world record, though a few matched or got close on the easier puzzles.
Humans are safe for now. That clearly won’t last.
Artificial Analysis upgrades its Intelligence Index to v4.1, shifting towards harder and more agentic tasks and consistently tracking time and money spent.
Opus 4.8 is the best available model by their metric in terms of result, slightly ahead of GPT-5.5, with a substantial gap down to everyone else. In exchange, GPT-5.5 was considerably cheaper and faster.
DeepSeek v4 cost only $0.04 per task for a score of 44, so it looks like a solid pick when you’re primarily looking for fast and cheap.
Fable 5 was substantially better than all of them, but is not currently available.
They also give us GDPval-AA v2 as part of this, which shows a similar pattern.
OpenAI gives us LifeSciBench, which is 750 expert-authored tasks spanning seven workflows and seven biological domains. They choose to compare GPT to Grok 4.3 and Gemini 3.1, so we have no idea if their score is any good.
Gemini can underperform on evals because sometimes it stops caring about the result and starts treating it like a puzzle or a consequence-free simulation. If Gemini thinks it is being tested on ethics it acts ethical, but in a free play space or roleplay with no consequences it (quite reasonably) acts less ethically instead. Very cool work. I buy that the uncertainty has to run in both directions.
It is very hard to get gains from specialization faster than the bitter lesson.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力