Claude Opus 4.7 发布:性能争议与自适应思考
Claude Opus 4.7 - A New Frontier, in Performance … and Drama
The kind of best AI model is here, Claude Opus 4.7, but it does bring with it a ton of controversy.
It's been out less than 24 hours, but I'm going to cover not only the bonanza of benchmarks that came out with it. Yes, including its score on my own simple bench. We'll hear Anthropic admit some unexpected flaws with the new model, as well as how in other areas they sabotage the capability of Opus. We'll look at the skill at which it falls behind some Gemini models, the areas in which it beats every other model, an industry first for Anthropic, but also why some users are furious at the company.
There are a bunch of cool Claude code and co-work upgrades, but also some strange downgrades in the default settings of Claude Opus. Plus, we have revelations about what OpenAI plan to do in response, along with a 9-year personal rivalry that comes to the fore.
That's a great list, Philip, but how good is it?
Well, kind of depends. Claude Opus 4.7 will think adaptively. In other words, if it thinks your task is easy, it will spend less time, quote, thinking about it. The benchmark I crafted, SimpleBench, contains a bunch of trick questions that basically require common sense to see through. Because Opus 4.7 seems to think these questions are easier than they actually are, it scores worse than Opus 4.6. But let me give you perhaps a more pertinent example for your workflow.
I actually regularly use the Claude series to update this benchmarks page on my web app, lmcouncil.ai. Without telling them to, all the previous Claude models would attach the open root of tooltip when a new model was added. Hover the benchmark score and get the tooltip. When Opus 4.7 added itself to the leaderboard, it was the first model to not bother to do that. I had to then instruct it to do so. Just anecdotal, of course, but the model definitely will decide how much time to spend on your task.
Now, across more industry standard benchmarks, here's what's remember. In almost every case, Opus 4.7 outperforms Opus 4.6, but underperforms Mythos preview, which of course you can't access. Whether that's coding, obscure knowledge, or being able to navigate your computer. Though again, interestingly, when it comes to agentic search, browsing the web to retrieve interesting bits of information, hard-to-find snippets, Opus 4.7 in browse comp underperforms Opus 4.6.
Indeed, on that benchmark, even Mythos preview underperforms GPT-4 5.4. We haven't even gotten to the real controversies involved in Opus 4.7's release, but even in terms of benchmarks, the picture isn't crystal clear. Here's another detail you might have noticed. Opus 4.7 underperforms both Opus 4.6 and Mythos preview when it comes to cybersecurity vulnerability reproduction. Seems bad until you realize that on page 48 of the system card, Anthropic say, "This underperformance is in line with our expectations.
During training, we experimented with efforts to differentially reduce these capabilities. They don't want Opus 4.7 to be too good at finding vulnerabilities. In certain measures of long context reasoning, reasoning through vast documents, Opus 4.7 is a clear improvement over Opus 4.6. In another one involving, for example, finding the fourth poem across 1 million tokens, it's a regression, even on the max setting. The lead creator of Claude code said, "We kept that in the system card for scientific honesty, but we're phasing out because it's built around stacking distractors to trick the model.
On certain benchmarks, like this generalized measure of knowledge work, we get direct comparisons between Opus 4.7 and competitive models like Gemini 3.1 Pro, with Opus 4.7 seeming to be the best at vanilla office work. This is presumably why on page three of the system card, the company famous for promising not to advance the rate of AI progress says that Opus 4.7 is on real-world professional tasks ahead of all generally available models.
But then, on other benchmarks, we don't get a comparison with, for example, the Gemini series. Take vision, where, based on resolution, it indeed is better at navigating really dense graphical interfaces. But when an external benchmark group gave it a comprehensive OCR test, testing model's ability to visually pass through documents, Opus 4.7 actually underperformed the dramatically cheaper Gemini 3 Flash. Yes, an improvement on Opus 4.6, but on average underperforming the model that's more than 10 times cheaper, Gemini 3 Flash.
On page 43 of the system card, on an aggregate measure of benchmarks, we see that Opus 4.7 is kind of in line with expected model progress, based on the previous performance of Claude models, with only Mythos as the slight exception. But here's where Anthropic admit, "Benchmark supply at the frontier remains a bottleneck." This is why trying to discuss a model's IQ or its progress toward superintelligence gets increasingly difficult.
There is no one universal metric of a model's ability. Depending on the data it's fed, it might do worse at an abstract pattern recognition benchmark like Arc-AGI-2. There, Claude 4.7 underperforms GPT-5.4 Pro. According to Valse AI, though, on vibe coding, building a web app from scratch, Opus 4.7 is the best, beating out GPT-5.4 on performance and speed, albeit not on cost. done in terms of benchmarks, so time for a metric that is almost impossible to game, and that is market share of generative AI website traffic.
Now, what you may notice is that both Gemini and Claude have, compared to this time last year, roughly 4x'd their market share. For the first time since the release of the original ChatGPT in November of 2022, the market share of OpenAI may fall below 50% this month. That seems great for Claude, right? Except there is one knock-on consequence of that. According to an Open AI memo leaked to The Verge, Open AI believe that Anthropic have made a strategic misstep in not acquiring enough compute.
And that, they say, is going to show up in product. Customers may already be feeling it through throttling, weaker availability, and a less reliable experience. That could explain why adaptive thinking is now mandatory if you want extended thinking from the model. In other words, you can't force it to think longer. You can encourage it to think about thinking longer, but you can't force the Claude models to always think longer, to always use more inference compute.
Not only that, when one AMD senior AI director said that Claude had been nerfed, that's 4.6 before even 4.7 came out, and she brought the receipts saying how the number of characters used for thinking had dropped by three quarters. There was less thinking, far more bailing out. The lead creator of Claude Code replied to that and said that medium effort was now the default. You have to actively set the effort at high or max.
This, it seems, is one of the big things that Open AI still has over Anthropic. It's almost like the runaway success of Claude has led to this one Achilles heel. Sam Altman implicitly joked about the rate limits of Claude or being forced to use worse models. One of the leads on Codex talked about how Codex is, in comparison, compute efficient, always up, never down. And even that comment was before the release of GPT 5.4 cyber, which could be their mythos tier model for cybersecurity, but we just don't know.
No metrics, no benchmarks. Only insiders get access. That brings me back to the leaked memo where Open AI's chief revenue officer also said this, "Anthropic's story is built on fear, restriction, and the idea that a small group of elites should control AI." Open AI's analysis shows that Anthropic's run rate is overstated by roughly 8 billion, so should be more like 22 billion. In other words, still behind OpenAI. More interesting to me is the personal tussle at the heart of this rivalry between OpenAI and Anthropic.
But, I want to save that for just a little later after I cover more of the essentials about 4.7 Opus. Before we leave the money front t
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力