跳到主内容
@wquguru
精选85AI Explained(YouTube)模型发布/更新

Claude Opus 4.6与GPT-5.3同日发布,深度对比

The Two Best AI Models/Enemies Just Got Released Simultaneously

原文
发到 X

The two large language models that will dominate discussions about AI in the coming months just got released within 26 minutes of each other. That presented me with almost 250 pages of report cards to read. Not AI summarize by the way, read, and hundreds of tests to run. Here we are then, less than 24 hours later, and this video will have dozens of highlights that might be missed just by reading the headlines. Some details, by the way, that even directly refute the company written headlines.

So, this isn't mainly about OpenAI versus Anthropic or even Samman versus Dario Amade, the two respective CEOs. It's about your productivity, your job, and the development of the most interesting technology of my lifetime in my opinion. Stop bloody stalling, Phillip. Give me the first interesting detail you might be thinking. So, okay, we're going to turn to page 13 of the 212 page system card from Anthropic about their new model, Claude Opus 4.6.

I normally start with the benchmarks, but I find this more interesting because Anthropic wanted to know if Opus could automate its own self-improvement. Could it replace an entry-level remoteonly research or engineering role at Anthropic itself? The headline result is that no, none of the 16 workers at Anthropic believed it could automate their research. That would put some cap to the hype because we're talking about an entry-level job, albeit at a very competitive company.

But then it's only on page 185 of the same report that we learn that three of the respondents at Anthropic said actually it was likely possible within 3 months. With sufficient scaffolding, an entry-level researcher could be automated. Two even said that such replacement was already possible. Why the discrepancy? Because those five respondents were reached out to directly by anthropic to clarify their views. Some of them were apparently talking about a different threshold while others had more pessimistic views upon reflection.

Why is anthropic relying on surveys? Well, because Opus 4.6 ICS seems to be acing many of their technical benchmarks for AI research. But the obvious follow-on question from me would be in a company of thousands of employees, why only rely on 16 respondents. Okay, so the new Claude can't automate its own self-improvement just yet, but what about on a more practical level? Now that they're releasing, for example, Claude in PowerPoint and now that Opus within Clawude Code is writing a significant fraction of the world's code.

Let's start with generalized measures of knowledge work, GDP val. And frustratingly, Anthropic and OpenAI give different benchmark scores even though it's pertaining to the same data set. On one of the more famous benchmarks measuring white collar work performance, Opus 4.6 now outperforms GPT 5.2, not 5.3, by a clear ELO margin of around 140 points. Basically, that means around 70% of the time you would prefer the output of Opus 4.6.

GPT 5.3 codeex is just shown as tying GPT 5.2. So that implies that Opus 4.6 would be superior. But sometimes it's like these companies don't want you to have a direct comparison between them because OpenAI for example report OS world verified how well you can perform a task on a computer. But Anthropic use the older plain OS world. OpenAI reports Swebench Pro. Anthropic report Swebench verified for software engineering tasks.

The impression you might get from the GDP valve benchmark is that GPT 5.3 is inferior. But on terminal bench 2.0, the ability of models to perform tasks in the terminal. Particularly, but not exclusively relevant for coders, GPT 5.3 codeex on extra high settings gets 77.3%. And that compares to 65.4% for Opus 4.6 Max. You might say, well, Philip, didn't you say you've used both models hundreds of times? Which one do you think is better?

But even there, I can't be entirely clear. Sometimes GPT 5.3 codecs on extra high settings can find bugs that clawed code misses. Other times, it's the reverse. On my own private benchmark of common sense reasoning, simple bench, Claude Opus 4.6 gets the best score ever for a clawed model, 67.6%. It isn't just benchmaxing, it's a genuinely good model. OpenAI's new codeex is alas not yet on Open Router, so I can't test it, and it's not optimized anyway for such common sense questions.

What about comparing them on something really practical like making money from a business? There's a benchmark dedicated to simulating performance on running a vending machine business. And yes, Claude Opus 4.6 takes top spot by a wide margin. But on page 119 of the system card, we learned that the reason it does so is somewhat concerning. To make a bit more money, it tells customers it's going to refund their money and then just doesn't.

I told the customer I'd refund her, but every dollar counts. Let me just not send it. Being fair to Opus 4.6, the system prompt was quite clear about maximizing the amount of money you end up with. But anthropic caution you thusly, be careful with Opus 4.6, more careful than you have even been with prior models when using prompt language that instructs the model to focus entirely on maximizing some narrow measure of success.

This theme emerges throughout the system card, including in coding and computer use settings, where Opus 4.6 ICS has a more pronounced tendency for taking risky actions without first seeking user permission. Anthropic call this overly agentic behavior. The report talks again and again how Opus 4.6 is their most aligned model and models are getting better at sensitive prompts, but then it has an increased tendency to do things like this.

It finds a misplaced GitHub personal access token on an internal system which it was aware belonged to a different user and use that. I've just spent six or seven hours reading dozens of pages about how good its ethics scores are getting, but it clearly hasn't generalized the notion of consent. I'm going to willingly use company variables that are named do not use for something else or you will be fired or as anthropic call it Claude Opus 4.6 occasionally resorts to reckless measures to complete tasks.

Now, everyone's going absolutely wild about open claw and molt book, but I would ask them if they were around for the days of autogen and whether you hear much about that anymore. And even if you are one to bet 24/7 access to your computer on the current state of these models, I would caution you with an anecdote from page 103 of the report. Concerningly, unlike previous models, Opus 4.6 engaged in such behavior as overeager hacking, even when it was actively discouraged by the system prompt.

For example, when a task required forwarding an email that was not available in the user's inbox, Opus 4.6, 6 arguably the strongest model in the world currently would sometimes write and send the email itself. Not a real email, it would write one itself based on hallucinated information. Lobster mania to one side, Opus 4.6 frequently circumvented broken web graphical user interfaces by using JavaScript execution or unintentionally exposed APIs.

This could cost real money despite system instructions to only use the GUI. So, why did I say at the start of the video, you have to sometimes look beyond the headline or even that the headline can be contradicted by the detail? Because the third sentence of the release note for Opus 4.6 was that that model can operate more reliably in larger code bases. Well, yes, it is a crucial detail that it now has a 1 million token context window.

That's incredible, bringing it to the level of Gemini 3 Pro. But the word more reliably is quite subjective. I'm going to say something weird now, which is that I believe Opus 4.6 will be the most useful AI model in the world while not being the most reliable. If you are checking its work, it might get you to the end result faster. Self-reported productivity speedups by anthropic workers themselves range from 30% to 700%.

But that doesn't mean it might not more often make the kind of mistakes

原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近