跳到主内容
@wquguru
精选80AI Explained(YouTube)模型发布/更新多源精选 ×14

GPT-6 Astra发布:超越Claude 5.1,OpenAI内部担忧

GPT 6 Astra, so good even OpenAI are worried

原文
发到 X

GPT6 is here at last and yes, it's yet another AI model. But GPT6 Astra deserves the new number and the new name. And if you know people who haven't checked in on AI in a few weeks or months, their opinions on models may be far more stale than they realize. Because as the title of this video says, Astra is actually a very impressive model, not just another row of great benchmark scores. It's not just marketing to say that Astra's ability to reason silently has many within OpenAI somewhat disconcerted.

The first half of the video then will cover why many, including me, feel it is so good, while the second half will cover those concerns. Of course, that will leave you to decide where the weight of the emphasis should go. But without further ado, why do people say it's so good? Well, you could say the story begins with their main rival, Anthropic, and their Clawude Fable 5.1 model released in the last few days. Anthropic could have picked any benchmarks to emphasize how good Fable 5.1 was, but naturally, they picked some of the hardest.

I'm going to try to convey how hard some of these benchmarks are in a moment. Things like Terminal Bench Science and Automation Bench. But what you need to know is these are the benchmarks that Anthropic chose. And those are the same benchmarks that GPT6 Astra beats Fable 5.1 at while doing so at a lower cost. We'll see this trend across dozens of benchmarks. But that will only persuade those of you who care most about which is the best model.

If you think all of the models are overhyped and a bit useless, we're going to have to zoom in on some of these benchmarks. Starting with, for example, Terminal Bench Science 0.1. The title is so boring. What does it actually mean to do well at this benchmark? Well, I use the model to explain the model's performance. Here's an example of what the model is tked to do with code to answer real scientific questions. It's given 25,000 brightness readings from three stars.

It's got to find a small repeating dip in those stars' brightness. Code up a program that will then work on six sets that the model can't see. The model sees thousands of rows that look a bit like these. Does getting better than expert performance for a few dollars sound a bit more impressive now? How about for climate science? Given almost three gigabytes of image data, thousands of daily pictures of 40 lakes in Greenland, deduce how and why they got drained across, for example, a 153day season.

And remember, these aren't guess until you get it right. Models are penalized if they come up with bad hypotheses. There's also somewhat messy data. You could say cloud and snow, as the model says, can obscure the exact amount of drainage on a given day. It needs to therefore connect evidence before and after. I could go on and on. 237 MRI images. Finding, measuring the main injury, labeling it on an image, penalized for wrong answers, and this is all for one benchmark question.

Okay, that's just one benchmark, but there are many benchmarks. What about all the others? Well, let me now turn to agents last exam. And yes, I will be talking about benchmarks that OpenAI didn't include. As you might have guessed, on this benchmark, Astra sets a new state-of-the-art performance level, beating Claude Fable 5 and doing so at far less cost. You might quickly say, "Well, wait, aren't the API prices the same?"

But Astra uses far fewer tokens. It's more token efficient, and that's why the costs can be lower. But that's a factoid for the enthusiasts. I'm more focused on the skeptics, those who say these benchmarks don't measure anything real. Well, UC Berkeley designed agents last exam to exactly measure real things. Thousands of tasks curated by experts with verifiable outcomes across 55 industries on economically valuable tasks.

To make it more vivid, here's an example where the model must master industrial machining software. It's got to use a tool used in real factories and the kind of data that real humans would get. The model, of course, doesn't know about the kind of hidden tests its answer will be graded on. For this question, as you can see, the grader samples 10,000 points on a hidden reference surface with critical points needing to land within.3 mm.

Without knowing that, Astro has to plan every cut and avoid crashes. with any detected collision or cut into the finished part, making the score zero. Or this question on molten plastic, given the same kind of inputs that an expert would be given. Or what about game design where it has to recreate maps, monsters, and battles in an existing older game with a separate vision model judging every floor of its output? I could go on and on.

These benchmarks were created to be exceptionally hard for models. Speaking of vision, let's take a quick look at Screen Spot Pro. I should barely bother to even say that yes, it's state-of-the-art inaccuracy around 92%. I dug into the original paper released just over a year ago, and you can see the kind of software that models are tested on. Adobe Premiere, Photoshop, Office 365, God help it. The focus here, by the way, isn't how impressive its output is with this software.

It's whether the model can literally navigate such complex screens accurately, interact with very complex graphical user interfaces. When we see the examples in a moment of the kind of things that Astra can create, which I'm sure you guys are going to go off and try yourselves, it's because Astra can navigate across the computer so well with such fine grain controls of the interface that it can output such impressive things.

These aren't fake benchmarks, in other words, and I haven't even gotten to the most impressive ones. Let's take a quick look at Frontier Math Tier 4. You can see for yourself it beats Fable at a far lower cost, but it's more the story of how just over a year ago this benchmark was created. It was the hardest tier of math questions created for Frontier Math by Epoch AI. One professor of mathematics said at the time, quote, I can barely solve some of these problems from my own field.

So, I hope AI entities can't get any of them right. I want them to score zero. End quote. At the time, models did score around zero. For example, Gemini 2.5 Pro and indeed even GPT5 released around exactly a year ago got 10 12 20%. Not only did GPT6 Astra score at peak 98% but and we'll come to this later in the video without reasoning it got 83%. No scratch pad, no chain of thought, just give me an answer, 83%. I know there will still be people watching going, "Oh yeah, but that's just like a fake benchmark.

What about like real stuff? Real mathematics. Okay, let's go to an assistant professor at Stanford. I'm not claiming to be familiar with this result, but he said GPT6 blows away a result by among others the legendary Terrence Tao blows away that result by a full order of magnitude about finding the distance between prime numbers. Not only that, it created a new simple starting point on the problem since the 1930s with this as the key analogy.

This is akin to finding a completely new chess opening after decades, subverting human traditions. Dare one say it's a bit like a move 37. In terms of other more real feeling results, how about model development? One open AAI researcher said this with Astra, a research integration cycle that used to take me at least a month of full-time work took just over a week with only part of his time spent steering it. The job itself was literally about improving our next model immediately.

I can feel the strong momentum of recursive self-improvement. Now, I also briefly had to go back to use 5.6 soul internally. That's when the gap hit me. You don't always notice intelligence going up when you get used to it. You notice it when it gets taken away. Notice, by the way, free tier users won't be getting Astra. It's being rolled out to API customers, pro subscribers first, and soon plus subscribers. But forget the enthusiasts for a secon

原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
OpenAI发布Astra模型获好评
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文
OpenAI发布GPT-6 Astra
Hacker News Best(web_list)原文
OpenAI发布GPT-6 Astra:ARC-AGI得分99.9%
Simon Willison 博客(RSS)原文

相似阅读

另一事件,读法相近