跳到主内容
@wquguru
精选70AI Explained(YouTube)模型发布/更新

OpenAI发布GPT-5.4,白领工作自动化能力提升

What the New ChatGPT 5.4 Means for the World

原文
发到 X

Just 48 hours after releasing GPT 5.3 instant, OpenAI have released GPT 5.4. So, either we are at the sharp end of the singularity or Sam Altman really wants the headlines to shift away from other matters. But, for real though, it is a big update and ignoring frontier AI developments does feel more costly for professionals than ever. Though, I do sympathize with those trying to keep up because this is possibly the murkiest the AI landscape has ever been.

We get vague posts on X with earliest access most often given to those who are going to praise models. We get leaked reports and then accusations and counter-accusations, prediction market manipulation, and seemingly endless new benchmarks created by the companies themselves amid genuine progress in AI. I pretty much gave up after reaching about 76 tabs on Chrome. So, I'm just going to give you the nine things I think people should know from yes, the last few hours as well as the last few days of discombobulating developments.

Because for me, GPT 5.4 is OpenAI's attempt at making a Codex or Claude Code, but for all white-collar professionals. The model was blind graded by experts against human outputs from across 44 white-collar occupations, selected by the way for their impact on GDP, hence the name of the benchmark GDP Val, and GPT 5.4 beats the human first attempt 70.8% of the time. If you include ties, it's 83% of the time. But, that headline leaves to one side catastrophic failures when the model makes mistake that a human wouldn't make, and the fact that these tasks drawn from the work performed across these 44 occupations are self-contained, digital, and not representative of the full range of tasks and the purpose of these occupations.

You may have also noticed one narrative violation that we should hastily move on from, which is that GPT-5.4 Pro, available only to the highest paying users, actually scores worse in this benchmark than GPT-5.4. But still, all those caveats aside, if we make the analogy with self-driving, we might have reached a satisfying level of safety, but we may have passed that moment where mile for mile or spreadsheet for spreadsheet, the autonomous agents like GPT-5.4 are better.

And as Waymo has shown, even 10x safety performance does not mean you get national or international adoption. Which is to say, white-collar work is still set to last through the end of this year, he says nervously. Now, if that benchmark result from GPT-5.4 sounds scary, well, then you might tell your boss about this benchmark before they fire you and leave things fully to a model like GPT-5.4. On questions that probe for hallucinations, according to this benchmark from Artificial Analysis, GPT-5.4 does well.

Not quite as well as GPT-5.3 Codex, but measured by overall accuracy, it is close to state of the art. But when GPT-5.4 gets things wrong, it is more likely than other models to BS an answer. You want to be lower on this chart, in other words, and it's up here at 89%, meaning when it gets things wrong, instead of admitting it doesn't know the answer, it will BS. Side note, this comes almost 3 years after Sam Altman said, and I reported him here as saying that by this time last year, we would no longer need to discuss hallucinations.

I can tell that that might have de-hyped you for a moment on GPT-5.4, so let's re-hype you for a moment with a demonstration of the ongoing, actually quite breathtaking progress in almost autonomous software development. In OpenAI's Codex, which is now available on Windows as well as on Mac, I asked, essentially, can you create an animated league table for Stockport County FC's progress during the season. And we get this, which looks pretty beautiful and indeed has a function where you can play through the season and see their league position change as it goes along.

Now, I did check the current position of Stockport and it was accurate on that, but of course you would need to check each of these results. The fact that it can one-shot this though, and think of all the web searches it had to do as well, really does show that OpenAI is trying to bring all the disparate capabilities of its different tools into one place. They say it incorporates the industry-leading coding capabilities of GPT-5.3 Codex while improving how the model works across tools, software environments, professional tasks.

And because for argument's sake, let's say AI can do 98% of the coding required for world-class software, and skeptics could say, "Well, there's always going to be that 2% or 1% of things that it can't do. That's why employment figures for developers are staying healthy." Well, the other consequence of that ability is that non-developers can now perform at a level that is almost as good as the very best. In other words, the lines between professions are blurring in the reflected light from blazing sand.

Across a multitude of benchmarks, models are getting better and better at computer use, too. And GPT-5.4 is particularly pronounced in progress in that direction. But stripping the jargon, what that means is that the loop is almost closed as the model can see and click with unprecedented accuracy to test its own outputs. I asked the model to create a timeline of Viking incursions into England in a given period, and it did extremely well.

But if you play the campaign, the graphics are missing something. So, you definitely wouldn't call it one-shot. Look at the longship. It's nice, but wait, is that London or Iona misplaced? Sheppey is not there. And so you can see the accuracy isn't incredible. I know that that's being pretty harsh, but it's just true to say. But that loop, it's almost closed. It will be able to accurately see some of these mistakes. Maybe it already can.

I'm running it in the background now to improve it. By the end of the video, I'll show you the new version. But what then, when the loop is closed and you get impeccable software, one shot? As OpenAI point out, what happens when you apply that to spreadsheets, documents, and presentations? I don't know about you, but this version on the left from GPT-5.4 is a lot nicer looking than this one from GPT-5.2. That's epic.

The hype train has completely left the station and the singularity is nearer than the end of the Premier League season. But mayhap we spoke too soon. For it seems we are in the spiky world of AI performance, where record-breaking performance in one domain, derived from the finest of distilled training data, does not guarantee that such data exists in another domain. All of which leads to very uneven progress. Let me give you some examples from the 35-page system card of GPT-5.4.

If we look at one internal machine learning benchmark from OpenAI testing a model's ability to solve machine learning tasks, then the progress is pretty dramatic. A doubling from around 12% with GPT-5.2 thinking to 23% with GPT-5.4 thinking. There's no GPT-5.3 Codex though on this chart. So let's turn to OpenAI proof Q&A, which I think is a fantastic benchmark from OpenAI. Albeit, of course, an internal one. This is a benchmark made from 20 internal research and engineering bottlenecks that were actually encountered at OpenAI.

Each one added at least a one-day delay to a major project. Solving these, in other words, would have led to millions of dollars of saving for OpenAI. The solutions to the bottlenecks, by the way, took at least a day again to solve. Tasks required models to diagnose and explain complex issues such as unexpected performance regressions, anomalous training metrics, or subtle bugs. GPT-5.4 thinking not only underperforms GPT-5.3 Codex, but also GPT-5.2 Codex and even GPT-5.2 thinking.

Again, this is the central debate at the moment in AI. By training models on various sets of specialized data, the big bet from people like Dario Amodei and Sam Altman is that by specializing in these specialisms, models will generalize across specialisms. That in the future, they might not require as much specialized t

原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近