OpenAI与Anthropic新模型将引发政府紧迫感,Arc-AGI-3基准揭示AI与人类差距
Two AI Models Set to “stir government urgency”, But Will This Challenge Undo Them?
Two exclusive reports indicate that there will be a qualitative leap in AI performance from each of the next AI models released by OpenAI and Anthropic. For OpenAI, this has meant shutting down the Sora app to spare computing resources for the new Spud model. And for Anthropic, makers of Claude, it has meant renewed interest from the Pentagon in reviving a deal to use Claude beyond the 6-month deadline recently set by the US government.
But this video will also dive into a brand new benchmark sure to be the talk of 2026, Arc-AGI-3. I've read the paper in full, but the headline result is that humans get 100% while the best AI models currently get less than half a percent. That might or might not be news to Jensen Huang, CEO of Nvidia, who this week said that artificial general intelligence has already been achieved. Let's start though with OpenAI's erotica bot, because the news there is that that erotic chatbot is not coming out.
After presumably spending billions optimizing it for engagement, it has been shelved apparently. According to the Financial Times, which is always my source for sexbot rumors, OpenAI needs the compute for Spud. It needs to drop its other side quests to focus on AGI deployment, rolling up everything it does into one super app. Jumping across to the information, apparently even OpenAI employees had complained that Sora, with its viral AI videos, was still just a drag on the company's computing resources.
In contrast, the Spud model is apparently very strong, according to Sam Altman. It will be ready in a few weeks, and it will really accelerate the economy. I know at this point some of you will be rolling your eyes going, "Oh, he would say that." But this article was a strange echo of one I had just read in Axios about Anthropic and their Claude series. Here's the key paragraph that I don't think many people have noticed.
About the new Claude series, Anthropic have warned the US government officials that the next big advance will supercharge both offensive and defensive cyber capabilities. It might even stir government agency to strike some kind of deal. If you're not familiar with the breakdown in the former deal that Anthropic had with the Pentagon, check out my recent series of videos. This article, by the way, from Axios further hints that the Pentagon might be rethinking that breach in the deal.
Apparently, even after the most public part of the spat with Anthropic and the Department of War, the key negotiating official said that they're still very close to an agreement. One detail I thought some of you may be interested in is how Anthropic might get back into good graces with the government. One of the advisers to Dario Amodei, the CEO of Anthropic, is Brad Gerstner. He's the architect of Trump accounts, which provide $1,000 to every newborn whose parents enroll.
The article goes on to speculate that what would happen if Anthropic agreed to part fund these accounts. Almost like a very early tentative step to universal equity. Anyway, everything so far might get you fairly hyped about the imminence of a new tier of AI. So what follows is hopefully going to add a bit of context to all of this. Because in the last 48 hours, we got Arc-AGI-3, the continuation of a benchmark series I've covered for years on this channel.
For the makers of the benchmark, as long as there is a gap between AI and human learning, we do not have AGI. I'm going to get to the paper in a moment, but my immediate response to that headline is that there is at least some gap between humans and say chimps on mathematical memory and speed. Chimps can actually track numbers that flash briefly on a screen better than humans can. So by that logic, humans aren't AGI either, a finding possibly substantiated by global events.
But that aside, the Arc-AGI-3 puzzles are genuinely really fun to try. Maybe I'm sad. Maybe I should say slightly fun to try. And I love that they manage to simultaneously test exploration, planning, memory, and goal setting. For example, nowhere on screen, and for models either, does it say that you need to move the icon around in order to manipulate the environment, or that the plus symbol, for example, will rotate the shape, the one in the bottom left corner, or even more importantly, that the goal is to make the bottom left shape resemble the shape up here.
None of those goals are stated, but like in real life, sometimes goals have to be inferred or self-produced. I did have some exclusive insight into Arc-AGI-3, but when so many benchmarks are narrow, to have a benchmark that does not rely on language, memorized knowledge, or cultural cues, one that is indeed abstract, the A in Arc stands for abstraction, is for me healthy for the field. But for you though, on the details, what does the terrible performance of current frontier models tell us about how 2026 will unfold?
Here are the highlights from the 21-page paper. First, what happened to Arc-AGI-1 and 2, which were saturated fairly recently by frontier AI models? They give a lovely graph in the paper showing the rapid improvement in the last 18 months on those benchmarks. If you're not familiar, instead of an interactive game, these were more static tests of pattern recognition on grids. Well, the authors make two big points. First, that inbuilt chain of thought reasoning that was publicly debuted with 01 preview back in September 2024, that genuinely allowed models to demonstrate a kind of fluid intelligence.
Think on the fly, combine patterns from their training data to reach an end goal. That's part of the explanation for the saturation of those previous benchmarks. The other part of the explanation is more intriguing. The authors say that because the public set and the private test set of those benchmarks were quite similar, then any model trained on an enormous amount of tasks representing a dense sampling of the task space automatically generated for this purpose, in other words, thousands of different guesses of what the private test set might be like, could essentially game the benchmark.
It's not direct memorization, it's a higher-level shortcut, a form of attack. They pointed out that models like Gemini 3, in their chain of thought, had giveaways that their training data may have resembled Arc-AGI-like tasks, either incidentally or intentionally. Going forward, the authors argue, private test sets need to be quite distinct and out of distribution compared to publicly available demonstration data. For Arc-AGI-3, the public test set is different and easier than the semi-private test set that is tested via API, and the fully private test set that is used for the competition.
It's a different distribution of tasks, far less gameable, even if AI labs are intentionally trying to mix Arc-AGI tasks into their training data. The goal of Arc-AGI-3 is to measure the residual gap between frontier AI and human-level AGI. When those potentially crazy new models come out in the coming days or weeks, what remaining gaps, deficiencies will they have compared to humans? And I would rather the authors of this paper say deficiencies rather than gaps, because in the methodology of the paper, we learn that AI performance is clamped at 100% or the human-derived baseline of 100%.
So even if one day they solve these interactive games more efficiently than humans, they'll only ever score 100%. In other words, AI getting 100% on this benchmark will not be taken as proof of AGI, or even strong evidence of it, because the most that models can get is 100%. But yet current performance is taken as evidence of them not being AGI. The benchmark is also turn-based, so the superior speed of AI models, or their better reflexes, is not counted in the test.
Nor is the relative cheapness of models counted for too much in benchmark scoring here, because scores on Arc-AGI-3 are not based on how many levels you solve, but on how many actions you took to solve those levels. Also, if a model takes more than five tim
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力