OpenAI发布GPT-6 Astra,官方宣称进入AGI时代
GPT-6-Astra Can Do Ambitious Things
GPT-6 Astra作为全新旗舰模型发布,官方直接对标AGI,能力跃迁明显,是行业核心事件,建议从业者关注其实际工程落地效果。
Astra is an excellent model. The jump from Sol to Astra is larger than the jump from Fable 5 to Fable 5.1. This is a big deal.
Astra 是一款出色的模型。从 Sol 到 Astra 的跨越,比从 Fable 5 到 Fable 5.1 的跨越更大。这是一件大事。
Astra is the best model for what one would broadly call ‘ambitious projects,’ and likely has the highest raw intelligence factor of any model. These are the largest jumps.
Astra 是处理人们广义上称为‘雄心勃勃的项目’的最佳模型,并且可能拥有任何模型中最高的原始智力因子。这些是最大的跨越。
It is amazing at doing things in 3D, or anything involving games. Astra also excels at computer use, and at subagent coordination.
它在执行 3D 任务或任何涉及游戏的事情方面表现出色。Astra 在计算机使用以及子代理协调方面也同样擅长。
Many benchmarks show dramatic jumps from all previous models. Where Astra is good, it can be in a league of its own.
许多基准测试显示,与所有先前模型相比都有显著的提升。在 Astra 擅长的领域,它可能处于一个独立的层级。
That does not mean Astra is in its own league across the board. Fable 5.1 is still a Claude. Astra is still a GPT. If you have a strong preference for one over the other, that still applies. For many purposes, especially involving back-and-forth discussions, Fable 5.1 is still my top choice. Fable remains my primary editor.
这并不意味着 Astra 在所有方面都处于独立层级。Fable 5.1 仍然是 Claude,Astra 仍然是 GPT。如果你强烈偏好其中某一个,这种偏好依然适用。对于许多目的,尤其是涉及来回讨论的场景,Fable 5.1 仍然是我的首选。Fable 仍然是我的主要编辑器。
If you want the best answer to your questions, you should ask both models.
如果你想要对你问题的最佳答案,你应该同时询问这两个模型。
Regular coding is getting less of a focus. Astra is not a quantum leap there, but of course it is very good and makes progress over Sol.
常规编码不再是重点。Astra 在这方面并非质的飞跃,但它当然非常出色,并相较于 Sol 取得了进步。
This is the first time a debate over whether a model ‘was AGI’ felt non-silly. I do not think it is AGI, and I would warn against the dangers of using that label prematurely, but I would not laugh at you for disagreeing.
这是第一次,关于某个模型是否‘是 AGI(通用人工智能)’的辩论感觉不再愚蠢。我认为它不是 AGI,我会警告过早使用该标签的危险,但如果你持不同意见,我也不会嘲笑你。
This is also a strange situation in that OpenAI has already soft announced that they have an internal model a level above Astra, as I will cover when I address Navier-Stokes.
这也是一种奇怪的情况,因为 OpenAI 已经软性宣布他们内部有一个比 Astra 高一个层级的模型,这将在我讨论 Navier-Stokes 时涵盖。
roon (OpenAI, September 3, 2026): I have not come close to discovering the limits of what Astra can do.
roon(OpenAI,2026年9月3日):我还没有接近发现 Astra 能力的极限。
I imagine it’ll be obsolete in the order of weeks somehow.
我想它在几周内的某个时候就会过时。
roon (OpenA, September 9, 2026): it didn’t even take a week.
roon(OpenA,2026年9月9日):甚至没花一周时间。
My recommendation is that you use both Fable 5.1 and Astra on your most difficult questions, and experiment to see which things each one does best for you.
我的建议是,在你最困难的问题上使用 Fable 5.1 和 Astra,并通过实验看看每个模型分别在哪方面对你最有帮助。
Astra, Self-Portrait
Astra,自画像
Table of Contents
目录
- Meanwhile.
- The Official Pitch.
- Our Price Cheap.
- Unnecessary Overstatement.
- Paced Rollout.
- Official Benchmarks.
- Other People’s Benchmarks.
- Thinking, Fast Without Slow.
- How Dare You, Sir.
- PoetryBench.
- In 3D.
- Time to Think.
- I’m Putting Together a Team.
- Reviews and Essays.
- Computer Use.
- Positive Reactions.
- AGI.
- Astra Can Do The Math.
- Astra Can Code.
- I Came to (Change the) Game.
- Astra Does Other Cool Things.
- Astra Never Quits Except When It Does.
- Negative Reactions.
- Stop It With the Hedging.
- Personality Clash.
- Revealed Preference.
- Dual Wielding.
- 与此同时。
- 官方宣传语。
- 我们的价格低廉。
- 不必要的夸大其词。
- 有节奏的发布。
- 官方基准测试。
- 其他人的基准测试。
- 思考:快而不慢。
- 你怎么敢,先生。
- PoetryBench。
- 在3D中。
- 思考的时间。
- 我正在组建团队。
- 评论与文章。
- 计算机使用。
- 积极反应。
- AGI(通用人工智能)。
- Astra能做数学题。
- Astra能写代码。
- 我来(改变)游戏规则了。
- Astra还做其他很酷的事情。
- Astra永不放弃,除非它真的放弃了。
- 消极反应。
- 别再含糊其辞了。
- 性格冲突。
- 显示偏好。
- 双持武器。
Meanwhile
与此同时
The backlog of things to discuss is not getting smaller, so I will take this moment to encourage everyone to read this excellent essay by Dario Amodei, We Must Pace the Frontier.
待讨论的事项积压并未减少,因此我借此机会鼓励大家阅读达里奥·阿莫代伊(Dario Amodei)的这篇优秀文章:《我们必须为前沿技术设定节奏》(We Must Pace the Frontier)。
This is the core thing he calls for, including a unilateral commitment:
这是他呼吁的核心内容,包括一项单方面承诺:
The steps are:
具体步骤如下:
- Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.
- Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.
- Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.
- 嵌入式评估员。每家前沿AI公司承诺向嵌入式的第三方评估员团队(如METR)提供持续、类似员工的访问权限。这些评估员的职责是验证对安全实践和承诺的遵守情况,报告事故,并帮助评估不仅限于已完成的AI模型,还包括训练管道和流程的对齐情况。这是任何节奏承诺可验证性的关键步骤,银行业已有先例,其中有时会将监管“监督员”嵌入员工之中。Anthropic目前正单方面承诺执行此步骤。我们旨在将此作为更广泛努力的一部分,以加倍投入我们的安全和对齐工作。
- 民主协调。民主国家内的前沿AI公司协调建立共同的安全标准以及限制无约束AI进步的速度。某些对设定节奏有影响力的协调形式在法律上具有挑战性,需要政府的支持。
- 全球协调。美国及其他民主政府试图与威权政府进行协调(在可能的范围内),同时认真对待验证合规性所带来的挑战。
I will have full coverage of that next week. Along with that essay, the current queue includes at least that, Navier-Stokes and Astra-2, Anthropic’s Misalignment Report, Anthropic’s Countering Misuse, a thinkpiece on personal AI and the law, a post called Claude Talk, and Fable 5.1 (and Astra?) Model Welfare.
我将在下周对此进行全面报道。除了那篇文章外,当前的队列至少包括:Navier-Stokes、Astra-2、Anthropic 的《对齐偏差报告》、Anthropic 的《遏制滥用》、一篇关于个人 AI 与法律的思想评论文章、一篇名为 Claude Talk 的文章,以及 Fable 5.1(和 Astra?)模型福利。
Okay, back to Astra.
好的,回到 Astra。
The Official Pitch
官方宣传口径
The pitch is that this is AGI.
其宣传点是:这是 AGI(通用人工智能)。
Greg Brockman (President, OpenAI): Welcome to the AGI era.
Greg Brockman(OpenAI 总裁):欢迎来到 AGI 时代。
Axios: OpenAI says Astra can lay out a printed circuit board in KiCad, build a 3D city scene in Unity, create an animated automobile transmission in FreeCAD and Blender, and fill out a tax-return draft from a W-2.
Axios:OpenAI 表示 Astra 能够在 KiCad 中绘制印刷电路板,在 Unity 中构建 3D 城市场景,在 FreeCAD 和 Blender 中创建动画汽车变速箱,并根据 W-2 表格填写税务申报草稿。
In scientific work, the model helped improve a mathematical result on gaps between prime numbers and set new marks on several biology, chemistry, medical and physics evaluations.
在科学研究方面,该模型帮助改进了关于素数间隔的数学结果,并在多项生物学、化学、医学和物理学评估中创下新纪录。
Jensen Huang also calls it AGI, but he calls everything AGI.
Jensen Huang 也称其为 AGI,但他称一切皆为 AGI。
The standard release video (3 min) has a stunning 125 million views on Twitter.
标准发布视频(3 分钟)在 Twitter 上获得了惊人的 1.25 亿次观看量。
Dean W. Ball (OpenAI): Astra is a remarkable piece of technology. Earlier agents often tried to dampen my ambitions—they’d push me to do “pilots” or “proofs of concept.” Then agents started meeting my ambitions.
Dean W. Ball(OpenAI):Astra 是一项非凡的技术。早期的代理往往试图压制我的雄心——它们会推动我去做“试点”或“概念验证”。然后,代理开始满足我的雄心。
Astra is the first agent that routinely raises my ambitions. I encourage you to try it!
Astra 是第一个经常提升我雄心的代理。我鼓励大家尝试一下!
I think it is a very very good writer, in a “big model smell” kind of way.
我认为它是一个非常非常优秀的写作者,带着一种“大模型气味”的风格。
roon (OpenAI): not to sound like a total shill but it’s the long weekend and all I want to do is make astrodynamical visualizations and stuff with Astra I have Astra psychosis
roon(OpenAI):听起来不像是在完全吹捧,但既然是长周末,我只想用 Astra 制作天体力学可视化之类的东西。我得了 Astra 精神病。
Tibo: Astra was probably our biggest competitive advantage while it wasn’t generally available.
Tibo:在 Astra 尚未普遍可用时,它可能是我们最大的竞争优势。
Since we’ve had it our productivity jumped so much that we shifted some of our plans 6 months ahead and will ship them at DevDay instead of mid next year.
自从我们拥有它以来,我们的生产力大幅提升,以至于我们将部分计划提前了 6 个月,并将在 DevDay 而不是明年年中交付它们。
Dominik Kundel says Astra can do all the things in Codex: Use all your apps, do more of the product thinking, excel at Blender, impress you without Max thinking and keep checking its work. He is very impressed.
Dominik Kundel 表示 Astra 能做 Codex 能做的所有事情:使用你所有的应用程序,进行更多的产品设计思考,精通 Blender,在不进行过度思考的情况下给你留下深刻印象,并持续检查其工作成果。他印象深刻。
Here is the pitch from Astra itself, according to Pangram:
以下是根据 Pangram 提供的 Astra 自身的宣传语:
Sam Altman (CEO OpenAI): GPT-6 Astra is here. We hope it will begin to enable a new generation of entrepreneurship, scientific discovery, and building.
Sam Altman(OpenAI CEO):GPT-6 Astra 来了。我们希望它能开启新一代创业、科学发现和建设的能力。
We believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more. It took us some extra time to ensure that we could meet the safety and alignment standards required for this capability level, but we think you’ll find it worth the wait.
我们相信,它是全球用于计算机操作、专业工作、科学计算、编码、网络安全等领域的最佳模型。为了确保能够达到这一能力水平所需的安全与对齐标准,我们多花了一些时间,但我们相信您会觉得等待是值得的。
It scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI 3, and 100% on ExploitBench.
它在 FrontierMath Tier 4 上得分 98%,在 ARC-AGI 3 上得分 99.9%,在 ExploitBench 上得分 100%。
Dr. Parik Patel, BA, CFA, ACCA Esq (top reply): I can’t wait to use this model to draft an email
Parik Patel 博士,BA,CFA,ACCA Esq(最高回复):我迫不及待想使用这个模型来起草一封电子邮件。
Mark Chen: GPT-6 Astra is here! This is a big moment for our research team – years of work on pretraining, reinforcement learning, and post-training have come together in our most capable and aligned model yet. It can build and test software, work across apps on your computer, and even help you take a crack at open scientific problems!
Mark Chen:GPT-6 Astra 来了!这对我们的研究团队来说是一个重要时刻——多年来在预训练、强化学习和后训练方面的工作汇聚于此,造就了目前我们最强大且对齐度最高的模型。它能够构建和测试软件,跨应用操作你的电脑,甚至能帮你尝试解决开放性的科学问题!
Capabilities that felt like grand challenges a few years ago have become tools people can actually use. One example is Computer Use – if you’ve tried this before and felt like it was too slow or not good enough, I encourage you to give it another shot. We’ve come a long way since Operator, and it “just works” now.
几年前还被视为重大挑战的能力,如今已成为人们真正可以使用的工具。一个例子就是 Computer Use(计算机使用)——如果你之前尝试过并觉得它太慢或不够好,我鼓励你再试一次。自 Operator 以来,我们已经取得了长足的进步,现在它“开箱即用”。
We’re also asking these systems to act on your behalf for more consequential work. Agents needs to stay aligned with your goals and values, think transparently, and respond to oversight even when tasks become difficult. We’ve made substantial progress on these behaviors in Astra, alongside stronger monitoring that can stop potentially unauthorized actions. That work is part of what makes this release possible.
我们还要求这些系统代表你执行更具影响力的任务。Agent 需要与你的目标和价值观保持一致,透明地思考,并在任务变得困难时响应监督。我们在 Astra 中在这些行为方面取得了实质性进展,同时增强了能够阻止潜在未经授权操作的监控机制。这些工作是此次发布得以实现的一部分原因。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力