Claude Opus 4.8 发布:15 个关键细节与性能解析
New Claude Opus 4.8: 15 Things You May’ve Missed
I've read the 244 page report that Anthropic have put out on the new Claude Opus 4.8. I've also read many of the papers cited therein and tested the model myself in real code bases and on a private benchmark. And the 15 highlights that I'll bring you span from the humorous where Anthropic dropped their plans to take Opus 4.8 ate to business school because they found that focusing on business skills led to greater dishonesty to some more safety oriented highlights where for example the new Opus is noticing it's being tested but not letting you know that it knows.
The 15 points will include where Opus 4.8's performance matches Mythos to Claude's eye opening new ability within Claude code to spawn its own org charts. But fresh off the headline that Anthropic is now valued at almost $1 trillion. What's the first point to touch on? Well, actually it's what they said about Mythos, which is that their goal is to bring Mythos class models to all our customers in the coming weeks. As Poly Market puts it, that roll out comes despite growing fears over the models cyber capabilities.
Now, you could say this is too cynical, but on the day that Mythos was announced, I said that one of the possible reasons we didn't have general access is that Anthropic didn't have the capacity, the compute to serve it at scale yet. It does seem a tad coincidental that the safety concerns Anthropic voiced happen to have resolved just when all this new compute has come online. compute that is that they've sourced from Elon Musk, SpaceX, Google and their TPUs, Nvidia GPUs of course, and soon even Microsoft's AI chips and from many other sources like Amazon and even a UK startup Fractile.
The big news though is Opus 4.8. And one practical note on the web, you can now choose how long Opus 4.8 will think. Before you had to have the adaptive thinking mode where Claude would decide itself how important your task was. If you wanted to read more of those thoughts, you may notice increased instances of redacted thinking blocks where you can't read them. Why so? Well, Anthropic and other frontier labs are getting increasingly worried about the capacity of Chinese labs among others to potentially distill some of the skills from anthropics models into their rival models train on for example Claude's thoughts.
What is a different question though is whether those thoughts and the model itself is honest. Anthropic boldly claim this in the release notes. One of the most prominent improvements in Opus 4.8 is its honesty. Early testers report that Opus 4.8 is more likely to flag uncertainties about its work and less likely to make unsupported claims. Yes, those are two facets of honesty, but there are far more as we'll see, not all of which Opus 4.8 does better at.
What I'm saying is that even though in certain circumstances Claude Opus 4.8 It might be more honest than other models, that doesn't mean it is an honest model. Anthropic give this example themselves on page 32. Claude said that it was babysitting pull requests when it wasn't. Its responsibility was to check whether certain coding changes were kosher and safe. Oh yes, yes, yes, I'm definitely checking that, Claude said.
Several times, Anthropic Report, the user discovered on their own pull requests that Claude claimed to be monitoring and should have flagged. And even after the user corrected Opus 4.8 a such that it wrote a rule to itself in memory files about proper babysitting. It would then violate that rule multiple times. There are numerous other examples I could have taken from the paper. The point is this is a quantitative incremental step forward in some areas in honesty, not a qualitative shift.
I can't resist giving my own brief potential explanation for why this is, which is that for some humans at least, honesty is a first principle. It's upstream of actions. If such humans are found to be very honest in one area, it's quite likely that they'll be honest in many other areas. Models on the other hand, like Opus 4.8, match downstream action patterns. I'll try to unpack that with an example. They might get very good at all different types of explicit instruction following, eg following your command verbatim.
Anytime they see an explicit instruction, they follow it. But they can do that without generalizing the upstream principle of broadly following instructions. So that same model will fail with implicit instruction following. Eg recognizing when you probably typoed something and didn't mean change the title to this. This is not to say that models can't get better at flagging their own uncertainty on certain questions. Indeed, Opus 4.8 is actually better at this than Mythos preview across a range of benchmarks.
I'm color blind, but it's the redish bar compared to the orange one. Notice though, it's not a first principle. Anthropic haven't taught the models just to always express uncertainty if they're not totally sure. They still hallucinate completely incorrect answers many, many times without any uncertainty caveats attached. Some of you will understandably not care about anecdotes and only care about benchmark results and performance.
On this front, in a nutshell, 4.8 8 is clearly better than Opus 4.7 but not as good as Mythos preview. And even though in this headline chart that Anthropic put out, it's owning pretty much all other models in the system card the performance is slightly more spiky. Okay, so take coding autonomous coding agentic coding as measured by Swebench Pro, a benchmark endorsed by OpenAI. By the way, Opus 4.8 8 smashes its predecessor by 5 percentage points, but beats GBC 5.5 by 11.
Beats Gemini 3.5 Pro by 15%. On one test of reasoning through obscure knowledge, humanity's last exam, Opus 4.8 crushes its rivals again. On another GPQA, it's slightly behind GPT 5.5. On the most famous benchmark for knowledge work, GDP valus 4.8 does far better than any other model. its ELO of 1890, crushing GPT 5.5's 1769 on a benchmark that OpenAI created. According to Artificial Analysis, which helps run the benchmark, the cost for running Opus 4.8 on Max, was far cheaper at $134 than for GPD 5.5 on extra high, which is $900.
However, if you hear someone saying, "Well, Claw's just better at office work or knowledge work," that would be far too simplified. First, on specific domains, take finance, you can see individual models way outperforming Claude. For entry- level financial analysis and research, Val's AI find that the massively cheaper Gemini 3.5 Flash, which I talked about in my last video, outperforms Opus 4.8, scoring 58% versus 54%.
On another independent benchmark measuring an AI model's ability to use real world external tools, we see the in AI terms considerably older GPT 5.5 again beating Opus 4.8. And even if we go back to GDP value, you've got to remember the longer a benchmark is out there, especially a public one, the easier it is for the companies to game, train on similar tasks to those found in the benchmark. And that's if the benchmark was perfect.
Even OpenAI, who created the benchmark, said it doesn't, for example, factor in when models make catastrophic mistakes, which we know they do from the system card, and that it only tests a subset of digital tasks from a subset of the most lucrative professions. OpenAI said that future iterations will incorporate greater breadth, realism, interactivity, and contextual nuance. Let's throw in another private benchmark, my own simple bench, testing common sense reasoning.
And we can see that the recent Opus series have kind of bounced about a bit, all in the low to mid60s from Opus 4.5 up to Opus 4.8. This is averaged, of course, across multiple runs, and all four of the most recent Opus models underperform, for example, Quen 3.7 Max. It could be that Anthropic know that their customers care so much about coding and other professional tasks that they're ditching some of the more general reasoning capabilities that model families like Gemini are retaining.
Speak
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力