GPT-5.5与DeepSeek V4同日发布,算力竞争白热化
GPT 5.5 Arrives, DeepSeek V4 Drops, and the Compute War Intensifies
GPT-5.5和DeepSeek V4同日发布,直接对标,影响数亿用户。做模型选型或关注竞争格局的同学必看,注意GPT-5.5在编码和幻觉上的权衡,以及DeepSeek的性价比冲击。
In the last 20 hours in AI, we have gotten two new models that could influence how a billion people use AI. In my mind, GPT-5.5 is OpenAI's all-out attempt to keep the AI crown from slipping to Anthropic. While today's DeepSeek V4 is China's answer to both. And in the swirl of headlines you are seeing today, you might have missed up to 50 data points that could affect how you work and how you use AI. So, I'm going to try and give you all of them plus select highlights from hours worth of interviews that I've watched with lab leaders.
You probably know me well enough to know that I've read the papers, too. So, we'll hear about OpenAI's updated estimate on the chances of recursive self-improvement. It was quite surprising. GPT-5.5's slight preference for men, which I'll explain. Mythos comparisons. And why the OpenAI president laughed at Anthropic's compute situation. For reference, I'll start with a focus on GPT-5.5, then do DeepSeek, and end by zooming out for the juiciest part of the overview.
For the brand new GPT-5.5, I did get early access, but there's no API access at the moment for anyone. So, almost all of the benchmark scores you're going to hear about are self-reported from OpenAI. I will say for me, testing out GPT-5.5 for days in the run-up to this release, it will become my daily driver, just about nudging out Opus 4.7. There's lots of caveats to that, though. As you can see, with GPT-5.5 underperforming both Opus 4.7 and of course Mythos preview on agentic coding, SWE-bench Pro.
Notice GPT-5.5 underperforms Opus 4.7 by around 6%, but Mythos preview by almost 20%. What you might not notice is that there's no entry for SWE-bench verified. And so, you might say, "Well, Philip, who cares about SWE-bench Pro then? What does it even mean that one row? Well, to Open AI, it seemingly means a lot because as Neel Chaudhuri points out, in February, Open AI told us to switch to Swebench Pro. That's the one it underperforms in because it's less contaminated than Swebench verified.
According to the Open AI blog post, we recommend Swebench Pro. You are probably going to go through a bit of a roller coaster in this video because if you look one row down at agentic terminal coding, you'll see GPT-5.5 way ahead. It's 82.7% score beating out Mythos Previews 82.0%. And so if you had just been feeling down about GPT-5.5's coding ability, there's another reminder I'll bring, which is that we've been talking about GPT-5.5, not even GPT-5.5 Pro, which is coming to the API very soon.
So while it's tempting to say that Mythos is absolutely mocking GPT-5.5, and let me know if I used that word correctly, we don't actually have an apples-to-apples comparison. The Mandate of Heaven is very much up for grabs. Okay, so now you're a bit confused. Let's look further. Let's look at Humanity's Last Exam, which is more of a arcane knowledge benchmark. Obscure academic domains combined with advanced reasoning.
Well, there GPT-5.5 is beaten by both Opus 4.7 and Mythos, as well as Gemini 3.1 Pro, by the way, without tools. But there's a caveat even to this because that involves a lot of general knowledge. It could well be that Open AI are at least slightly deemphasizing such general knowledge to make the model more efficient and cheaper. One of the top researchers at Open AI who I've been quoting for years, Noam Brown, said, "What matters is intelligence per token or per dollar.
After all, if you spend more, you do go up in benchmark score." Or in fancier language, intelligence is a function of inference compute. That being the case, if GPT-5.5 can work well in the domains you care about and use fewer tokens to get the answers you care about, then you may just frankly not care about Humanity's Last Exam. In one famous test of pattern recognition, ARC AGI 2, you'll see that GPT-5.5 on all settings beats out the Claude Opus series 4.6 and 4.7.
Not only achieving higher scores, but for much lower cost. Just one benchmark, of course, but we have to increasingly focus on performance per dollar these days. And on that front, DeepSeek will definitely want a word. Because holy moly, I'll get to them later, but DeepSeek V4 Pro got 61.2% in my own private benchmark, Simple Bench. It asks spatio-temporal questions that you need common sense to see through the tricks of.
But to get within 1 or 2% of Opus 4.7, I wasn't expecting that. At an absolute fraction of the cost, by the way. Again, no GPT-5.5 score because no API access. What about those frantic headlines about Mythos being able to hack into virtually any system? I think a lot of that was overblown, and some of that could be achieved by much smaller models. But nevertheless, skipping to page 33 of the system card, you can see that one external institute, the UK AI Security Institute, judges that GPT-5.5 is the strongest-performing model overall on their narrow cyber tasks, albeit within the margin of error.
This section was notably vague with a headline score implying that 5.5 was better than Mythos, i.e. better than any other model they've tested. But then on their end-to-end cyber range task, 5.5 was able to complete a task in full on one out of 10 attempts, a 32-step corporate network attack simulation. One would take an expert 20 hours. Mythos, it seems though, could do it in three out of 10 attempts. As you can see, direct comparison is hard, but 5.5 does at least seem to be in the ballpark of Mythos's capabilities.
In other words, small-scale enterprise networks with weak security posture and a lack of defensive tooling could be vulnerable to autonomous end-to-end cyberattack capability via 5.5. Of course, there are additional safeguards put on top of 5.5 to prevent that happening. But given that the world's top bankers and CEOs have gotten together to discuss the risk of Mythos releasing a comparable model without nearly as much cybersecurity fanfare does indicate a rather profound difference of perspective.
Here's Sam Altman on the Mythos marketing. There are people in the world who for a long time have wanted to keep AI in the hands of a smaller group of people. Um you can justify that in a lot of different ways and some of it's real. There are going to be legitimate safety concerns. Um but I expect but if what you want is like we need control of AI just us cuz we're the trustworthy people, I think that the fear-based marketing is probably the most effective way to justify that.
Um that doesn't mean it's not legitimate in some cases. Uh but it is you know, clearly incredible marketing to say, "We have built a bomb. We're about to drop it on your head. We will sell you a bomb shelter for $100 million. You need to like run across all your stuff but only if we like pick you as a customer." Well, there's another way that we could compare GPT 5.5 with Mythos and that's to look at hallucinations. Ask the models a bunch of obscure knowledge questions and see how many they get right and just as importantly, how many of the ones they get wrong they admit to not knowing.
The headline score looks amazing. GPT 5.5 gets the most right, 57% versus Opus 4.6 and 4.7's 46% and I know Mythos isn't on there but I'll get to that. However, as we've learned on this channel, headlines can be misleading. Look at the hallucination rate. That's the questions it gets wrong and should have said, "I don't know." instead of hallucinating, fabricating an answer. Whoa there, GPT 5.5 at 86% hallucinating 86% of the questions it got wrong rather than saying, "I don't know."
Opus 4.7 on max, just 36%. Okay then, well let's focus on the net rate, the overall rate factoring in both correct and incorrect. We have a slight win for Opus 4.7 over GPT 5.5, 26 versus 20. But, here's where Mythos comes in. Because, buried fairly deep in the Opus 4.7 system card on page 126, we get a comparison between Opus 4.6, Opus 4.7, and Mythos. We can then compare Mythos with GPT 5.5 on extra high. Notice how Mythos gets way more correct, 71%.
Still hallucinating, of course,
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力