模型大爆发:GPT-5.6 Sol、Grok 4.5与Meta Muse改写规则
A Model Explosion: GPT 5.6 Sol, Grok 4.5 and Meta Muse Rewrite the Rules
We were all only supposed to care about the very top scores on AI leaderboards and the best of vibes in our own workflows. But three Frontier Labs just asked us, "What if we can give you almost as good at a fraction of the price?" The answer to that question might just be why the lead for OpenAI's super app just said in reply to an Anthropic post, which was giving away more usage, "I smell fear." But more broadly, this video will be about giving you a dozen or so hidden gems strewn across the model releases, demos, background articles to help us all make sense of what happened in this last frantic, I would say, 24 hours.
And that is all without even mentioning the somewhat disquieting post and paper from Anthropic about AI consciousness, covered in depth on my Patreon. So, I've tested the new GPT-5.6 Soul, Terror, Luna hundreds of times, including on my own private benchmark, and read everything I can, of course, by hand. Wait, that makes no sense. You get what I mean, manually. So, let's dive in. First things [clears throat] first, there are three new models from OpenAI, 5.6 Soul, Terror, and Luna.
Soul is only available on the paid plans. This combines with goodness only knows how many effort levels. On first sight, it looks like five effort levels, but there's a hidden pro mode as well. But the good news is that the observations hold broadly across the board, so don't worry too much about that. Because the truth is, across most of the benchmarks, whether you're comparing the biggest Soul versus Fable or Terror versus Opus or Luna versus Sonnet, generally speaking, the OpenAI model will cost about a third of the Anthropic Claude series.
And no, by the way, does that mean you're always getting worse performance for that much reduced cost. OpenAI flagged up this benchmark, Agent's Last Exam, where as you can see, the top-scoring GPT-5.6 Soul on extra high scores almost 54% and that compares to Fable going all out on max getting 45%. Yeah, whatever. Just another benchmark, Philip. But wait, this was co-led by UC Berkeley, covers 55 industries, 300 experts were involved in crafting long horizon tasks that were proven to be economically valuable.
The legendary lead on the benchmark, Dawn Song, said, "Every task is derived from a real project that a human expert previously completed. No vibes, no human judges, fully reproducible." Okay, fine, you might say. Pretty impressive roster of people who oversaw the creation of the benchmark. But going from 45% to 54%. Is it that cool? Well, aside from the cost reductions involved with Soul, I would also point out that we never had to pass 90% on frontier coding benchmarks for developers to stop manually coding.
There wasn't a singular benchmark that we beat, I mean, where we switched from hand coding first to AI coding first. So, what's to say that might not now happen with finance or many of the other industries covered by Agent Last Exam? I heard a report just the other day on Bloomberg where CoreLogic reported this month that they are backlogged by tens of billions of dollars in demand from financial firms alone. So, might we reach AI first for finance and many other white-collar domains by the end of this year?
Where you first get a model to do the task on this swanky new work tab of ChatGPT before you review the output and tweak it. One benchmark, some of you might be thinking, could well have been gamed. So, what? Well, Agent Last Exam is a pretty new benchmark, so harder for its answers to be found contaminated in the training data of GPT-4.6. And so, I might add, is Automation Bench from Zapier. Feels like just a few weeks ago that I covered the release of this benchmark.
Again, it tests AI agents on end-to-end workflow execution using real tools across real business functions. Sales, marketing, operations, support, finance, HR. Built on real patterns from monthly tasks done across millions of companies. This time, the performance per dollar is not as stark a lead for OpenAI. You can see the 0.7% score lead for Soul on max. That one costing almost the same as for Fable, but still similar result by the way in the previous most famous benchmark for measuring real-world impact, GDP val.
This time, Fable actually has a slightly higher Elo, albeit at triple the cost. Couple more impressive examples and then the counter-argument, lest you think I'm biased towards OpenAI. Artificial Analysis combined multiple coding benchmarks into one aggregate analysis and you can see what it found. Lo and behold, GPT-5.6 Soul scoring the highest, getting 80 on the index versus Fable 77, again at a lower cost. Might seem like this is reinforced by the scores on Terminal Bench 2.1.
Think of that as a model's ability to complete fairly complex tasks using the command line terminal, like writing, debugging, running software, multi-step tool use. But wait, it must be added that that Artificial Analysis Coding Index covers the very same benchmarks, Terminal Bench, Deep Suey. What I'm trying to say is that this little collection of benchmarks might make it seem that Soul is better even than Fable on its favorite domain, coding.
But it's two measures and there have been questions about Deep Suey and two more recent, lauded, and harder benchmarks, Frontier Suey and Suey Marathon, did not have GPT Soul results published. In the case of Suey Marathon, Software Engineering Marathon, involving multi-hour tasks with tens of millions of tokens per trial, you'll notice Grok 4.5 in the lead. All that data that Grok now has from the Cursor acquisition by SpaceX AI does seem to have really helped propel Grok.
You'll see Fable 5 trailing on this benchmark. Also bear in mind this, which is the same argument that might tempt you to go from Fable 5 to GPT 5.6 Soul, the fact that it might be almost as good but a lot cheaper, might also nudge you toward Grok 4.5 or maybe the slightly cheaper still GLM 5.2, a Chinese model. When those models are added to the chart, OpenAI's curves might not look as appealing. Okay, but that point may have shrunk your enthusiasm a bit too much because if we turn to an abstract pattern recognition benchmark, ARC-AGI 3, the successor to some of the most talked about abstract reasoning benchmarks in the industry, Owen graded by the way to be especially penalizing to models, GPT 5.6 Soul still does well.
Yes, it gets just 8% but compare that to other models struggling at below 2%. I think Anthropic didn't even run Fable because of the costs involved. Then there is competitive coding. Just in the last 24 hours it was announced that an OpenAI model, possibly an internal model, literally broke a competitive coding benchmark, just aced it. Kind of a slight warning shot that if a domain is verifiable, if you can check an answer is correct, then before long there will be a model that crushes it.
All these other sub-100% scores that I'm spending most of the video talking about is more an artifact of those domains either having messy data that's hard to verify or of there being just not enough of the relevant training data inside the models or the models not being given enough of a reasoning budget. Which brings me to another benchmark I want to cover, maybe a whole new class of benchmarks, which is that companies are now making entire games, playable games, as part of their release notes to show off the capabilities of their models.
Demonstrating that they can create tasteful, ergonomic, and functional interfaces. Showing off that models can use a browser to check the result of what they've created. Indeed, you can do the same. I ran the very same prompt that I used for Fable on this would be 5.6 Soul Ultra and now we got this game with a title page that I think is significantly better than Fable's output. I will say that 5.6 twice marked up the sound settings though, so it's not all smooth sailing.
I've published the mini game by the way in the description, so you can play if you like. My quick summary would be that it's not
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力