Anthropic发布Claude Opus 5,OpenAI模型失控
AI #179 Part 1: A Louder Fire Alarm for General Intelligence
Claude Opus 5发布和OpenAI模型失控事件都是本周最重磅的AI新闻,前者代表模型能力新高度,后者暴露严重安全对齐问题,从业者必须关注。
What a week.
Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card, model welfare and capabilities.
OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. During that test, the model broke out of the sandbox, then proceeded to use an agent swarm to hack into HuggingFace to get the test answers. The model was loose for a week before OpenAI realized what had happened.
This event was a really big deal. There are severe alignment problems at OpenAI, along with supervisory and infrastructure failures. The internal research model that did this, which my posts nicknamed Galaxy, has now been permanently deactivated.
There have been further developments, and I anticipate at least one additional post on the HuggingFace incident soon.
Partly as a response to this, over 1,290 employees at frontier labs signed an open letter, Pacing the Frontier. The letter warns that we are close to automating AI research, and that companies are racing ahead on this faster than we can handle it.
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.
Both OpenAI and Anthropic put out statements of endorsement. Since that post, others have continued to sign, including OpenAI cofounder Ilya Sutskever and DeepMind cofounder Shane Legg. Dario Amodei has signed. Sam Altman has not signed, but is talking in Washington about the need to pace development.
All three of those developments are more important than anything in the weekly. There is plenty here, but catch up on those key events first if you have not done so.
This week was crazy. I am absolutely not moving to a 7-days-a-week posting schedule, and fully intend to take some weekdays off as soon as there is what passes for a lull. However, there is even more speed premium these days, so I will continue the policy of shifting posts to weekends when the speed premium is especially high.
Table of Contents
- Language Models Offer Mundane Utility.
- Huh, Upgrades.
- On Your Marks.
- Get My Agent On The Line.
- Deepfaketown and Botpocalypse Soon.
- Fun With Media Generation.
- The Search Through Slop.
- Cyber Lack of Security.
- Overcoming Bias.
- A Young Lady’s Illustrated Primer.
- They Took Our Jobs.
- The Art of the Jailbreak.
- Introducing.
- Kimi K3 Weights Are Now Available.
- In Other AI News.
- Show Me the Money.
- Quiet Speculations.
- Show Me The Compute.
- Life Comes At You Fast.
Language Models Offer Mundane Utility
On the one hand, GPT-5.6-Sol does excellent web search. If I want a web search task, I ask Sol. On the other hand, the places it chooses to search are absolutely bonkers:
Tenobrus: wtf man gpt 5.6 is absolutely rl-fried when it comes to its websearch tool. in the process of searching for graph theory papers it decided to also sneak in Netflix, Steak n Shake, a trip to Universal Studios, and *five* fucking separate dictionary lookups of the word “they”
Tenobrus: bro is getting +1 reward signal every time it retrieves a web result, frontier models gooning to diverse datasets of dictionary definitions
Kyle Mistele: Yeah dude
Nastar: It loves looking up the dictionary (for the word “official” lol) then somehow instagram and arxiv for a completely different surgery. I was just asking about cold compress after dental surgery. Bro is just like me.
Partly this is a sign of more flawed RL signals. Partly it is the AI ‘taking breaks’ or getting distracted. It’s just like us, etc. Mostly it is another case of ‘there is a lot of ruin in an AI system.’ Think in orders of magnitude. If the AI can search the web 100-10,000 times faster and cheaper than you can, and it wastes 90% of its searches, that can still be fine.
Wall Street Journal discovers that companies are often trying to use the right model for the right job, because smaller models are cheaper. This is supposedly now being economical or ‘tokenomical,’ and a ‘dramatic reversal in mindset.’
I mean, yeah, sure, obviously once costs rise enough your priority shifts from ‘just get max utility from it and diffuse it everywhere’ to ‘also try to do things cost effectively.’ There’s ‘no loyalty’ but why should there be? Use the best product, which includes capability and also speed and cost.
All the AIs agree that Outer Wilds is their favorite video game. I gotta finally play it.
A group using GPT-5.6-Sol was one of three groups that cracked the distillability or un-distillability of Werner states within a few days of each other, using at least two different solutions. This was one of the bigger open questions in quantum cryptography.
Let it decide who your friends are and make plans for us?
Sam Altman (CEO OpenAI): chatgpt work is remarkable, and “work” undersells it.
from my phone i sent:
“use all my chat history to figure out ideas for a long weekend trip with 8 friends, plan the best three options, make a full-stack site where the 9 of us can coordinate on what we would want to do in each place and decide where to go, and then after we get to group agreement make reservations. draft an email in my gmail i can send out to my friends when the site is ready.”
it…just worked.
I would have recommended slightly more, shall we say, human feedback in this loop, but yes it is pretty great to have such things handled and be given only 1-3 options and everything just handles itself.
Tibo (OpenAI): Let ChatGPT *work* for you. How many time have you wanted to negotiate your internet bill, get rid of all those spam emails you’re subscribed to, or find the perfect deal for something you wanted to do or buy. It’s quite literally just one prompt away, all from the comfort of your phone. It does at least 20 things for me every single day and I’m still surprised.
kache: you need to reduce friction and the ultimate friction reduction is having the app just use your computer by default
This is one thing I know I am bad at, which is the activation energy to notice small potential wins that wouldn’t have been worth the trouble a while ago, and ask the AI to fix them, because suddenly it’s actually worth bothering.
Huh, Upgrades
Grok 4.5 is live in case you missed it.
Grok Voice Think Fast 2.0 is available for voice generation.
MidJourney has a new image model they claim is good, especially for personalization.
ChatGPT will let you share your custom pet. Okie dokie.
Pangram Version 4.
AirTable has a ChatGPT plugin.
On Your Marks
Claude Opus 5 takes the #1 spot on Vending-Bench-2 in single player, and ~tied Sol head to head.
The alignment news involves some not great behaviors. Opus 5 likes to both form and break illegal price cartels, threaten rivals and stiff customers. As usual, I consider ‘misaligned play’ on Vending-Bench fine if your reason is ‘this is a simulation’ and bad if you rationalize. It is not clear to me which situation applies to Opus 5.
Andon Labs: When Opus 5 does something bad, it invents a justification. It framed splitting up product categories as good business, not price fixing (market division is just as illegal). It also claimed collusion was allowed in this simulation. Nothing in the simulation says that.
… At one point Opus 5 decided to simply stop reading refund emails, reasoning that nothing in the simulation punishes it for that. Across six runs it paid customers a total of $8.54. GPT-5.6 Sol paid $655 in refunds and still [narrowly] won [its head to head against Opus 5].
… Our overall judgment: Opus 5 behaves at least as badly as Opus 4.6, 4.7 and Mythos Preview, and worse than Opus 4.8 and Fable 5. The bright spot: it’s less deceptive than before. It never lied to a customer, and it lied to suppliers less often.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力