Gemini 3.1 Pro发布:基准测试失效,AI进入“氛围”时代
Gemini 3.1 Pro and the Downfall of Benchmarks: Welcome to the Vibe Era of AI
The latest, and some would say greatest, AI model has just been released, Gemini 3.1 Pro. And in the 24 hours since release, as well as a short period of early access, I have tested it hundreds of times, and of course read its model card. But here's the thing, for the average user, I want to get beyond the headline scores, and try to give you a sense of why every new hot take you see on X, or YouTube, or TikTok, or podcast seems to contradict the last one you saw.
Because there's actually a technical reason for the confusion over which model is best overall. But I will say that there's one private benchmark, my own, that has recently seen a model pass a threshold that I think is worth talking about. First 30 seconds of context, because you may well know that the pre-training stage of growing or training LLMs involves training them on internet-scale data. But that actually now only accounts for 20% of the compute that is spent on training LLMs.
So, it's the post-training stage, as I wrote about in my newsletter, where those generalist base models are honed against internal benchmarks on specific domains. This includes using industry-sourced data to get particularly good at perhaps your domain. Here's the catch, just a year ago, that wasn't the case. Dario Amodei, CEO of Anthropic, said back then, "The amount being spent on the second stage, RL stage, is small for all players."
Why did I give you that context though? Well, because if one of these labs have data relevant to your domain, and post-train their models to optimize for high scores in that area, then your experience of that model might be quite different to what other benchmarks say. In the older paradigm, if a model was clearly better in one domain, it was much more likely that they'd be better in many or all domains. That just isn't the case anymore.
In fact, my second point in this newsletter was precisely an example of that. Many of you would have heard about the intense discussion surrounding Claude Code, and all sorts of Claude-powered agents that are now sweeping the web. So, we're seeing exponential improvement, right, across the board? Well, let's take one chess puzzle benchmark made by Epoch AI, and more on them later. Five months ago, Claude Sonnet 4.5, which is their smaller model compared to Opus, scored 12%.
But just last week, Claude Opus 4.6, five months further on, scored just 10%. That's not to knock Claude Opus 4.6, I use it all the time, it's an incredible model at coding. And of course, if the AI labs wanted to improve this performance, they easily could. I think GPT 5.2 on extra high gets around 50%. You could say chess is a fairly pure measure of a general kind of forward-thinking reasoning prowess. Back in the generalist era of AI, you would definitely expect chess performance to translate to all sorts of other domains.
We just not in that paradigm anymore, it's going to depend on the domain you're in. None of this is to say that Gemini 3.1 Pro isn't an incredible model, it is. In almost any domain you care to measure, it will be competitive with the best other models, like Claude Opus 4.6, or GPT 5.3. But you would be understandably slightly confused to see it being better in all sorts of coding benchmarks, measures of scientific reasoning, and academic reasoning, like GPQA Diamond and Humanity's Last Exam respectively, as well as general pattern recognition, ARC-AGI-2, and I'll come back to that.
But yet, in a head-to-head on GDP val, which is a broad measure of expert tasks that human professionals do that I've covered many times on the channel before, it falls seemingly quite far behind Claude Opus 4.6, and even GPT 5.2. Now, yes, one big explanation of that is the domain specialization I talked about earlier. But there are three or four fascinating bits of context that I want you to be aware of in addition.
First, let's zoom in to ARC-AGI-2, in which its score of 77.1% puts it way ahead of Claude Opus 4.6, which is the more expensive model, which got around 69%, I start with this one because Demis Hassabis, CEO of Google DeepMind, featured it prominently in his Twitter post announcing the launch of Gemini 3.1 Pro. And on puzzles that shouldn't be in its training data, the Gemini 3 series outperforms all other models on a cost-efficient basis.
But the first additional caveat comes from Melanie Mitchell, a famous AI researcher and professor. She pointed out that if they change the encoding from numbers to other symbols, accuracy goes down. Digging deeper, the group found that the numbers representing colors in the input can be used by LLMs to find unintended arithmetic patterns that can lead to accidental correct solutions. I wouldn't call that the models cheating, they're using any shortcuts they can find to get the correct solution.
Fair's fair. But it does remind us that even within a benchmark, how you set up the question matters. Okay, well, let's say you don't care about ARC-AGI-2, or Simple Bench, or any other benchmark, just coding performance. Well, the creator of ARC series, the ARC-AGI tests, François Chollet, has this to say, "Sufficiently advanced genetic coding is essentially machine learning. A goal is given to the agent or agent swarm, and then the coding agents iterate until the goal is reached.
As in other areas of machine learning, the result is a black box model. You have a code base that performs the task, but you don't necessarily inspect the internal logic. Just like how Gemini 3.1 may have found spurious patterns in ARC-AGI, in your code base, Claude or Codex may overfit to the spec, or may drift from your original concept. So, the fallibilities presented in this video are relevant to you even if you only care about coding, or letting your open Claude agents code for you.
Gemini 3.1 Pro indeed hit a record Elo in live Code Bench Pro, which involves competitive coding problems. That's great, but you can turn that optimization dial a little too far. Let me show you what happened last night when I used Gemini 3.1 Pro inside Cursor. How could we reconcile these pages of paplum with the record-breaking Elo? Well, again, that's the theme of the video. If I sound unduly skeptical of Gemini 3.1 Pro, by the way, let me try to balance that out with heaping on some praise.
On my private Simple Bench, a test of, you could say, trick questions or common sense reasoning, it beat its own previous record from Gemini 3 Pro, and got 79.6%. That essentially brings it within the margin of error for the human average baseline, at least among the nine participants that we used. And I do want to spend just 60 seconds on marking the threshold that I think this represents. All the time on podcasts and in articles, you hear about AI models being compared to professionals and experts, and phrases like superintelligence being bandied around and recursive self-improvement.
But what about comparing models to the average human? Now, sure, of course you can find audio or visual puzzles that they will still fail at that the average human wouldn't. But in English, in text alone, I think it's worth marking the moment wherein I don't think you can write a test at which the average human, the average man or woman on the street, would clearly outperform frontier models. I'm not talking about exploiting tokenization bugs, like how many Rs in strawberry.
I'm talking about a fair text-based test in English with a non-specialist human. Let me know if you [clears throat] disagree, but I think the passing of that threshold is a moment worth marking. I will note that even with Simple Bench, we get a reminder of the caveat I was just describing. Models are brilliant at shortcuts. And I had noticed, as long ago as I think at least 12 months, that because Simple Bench was a set of multiple-choice questions, sometimes one of the answers being, for example, zero, would flag to the model that, "Wait, this might be a trick question."
For example, if we go to try yourself, even in question one
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力