OpenAI发布GPT-6 Astra:旗舰模型与争议性评测
[AINews] GPT-6 Astra: OpenAI’s biggest LLM launch of all time
这是OpenAI年度级旗舰发布,不仅涉及全新模型架构与定价策略,更引发了关于模型真实能力边界、成本效益比及安全对齐性的深度行业辩论,极具讨论价值。
The launch is barely 9 hours old, and with 36M views and 164K likes, already is OpenAI’s most successful launch since Sora and certainly GPT-4 or GPT-5.
发布才刚刚过去9小时,凭借3600万次观看和16.4万个赞,这已经是OpenAI自Sora以来最成功的发布,也无疑是GPT-4或GPT-5以来最成功的一次。
You’ll recall we’ve previously observed that Anthropic tends to far outclass OpenAI in launch popularity. For the first time in their mutual history, OpenAI has turned the tables.
大家可能还记得,我们此前曾观察到Anthropic在发布热度上往往大幅领先OpenAI。这是它们相互竞争的历史中,OpenAI首次扭转了局面。
You can read our initial impressions here and we will update with more coverage soon, just stay subscribed.
你可以在这里阅读我们的初步印象,我们将很快带来更多报道,请保持订阅。
Overall a very welcome answer to Anthropic’s Fable and Opus progress.
总体而言,这是对Anthropic的Fable和Opus进展的一个非常令人欢迎的回应。
Your move, SpaceXAI and Google DeepMind.
轮到你们了,SpaceXAI和Google DeepMind。
AI News for 9/2/2026-9/3/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
2026年9月2日至9月3日AI新闻。我们检查了12个Reddit子版块、544条Twitter帖子,没有进一步的Discord讨论。AINews的网站允许你搜索所有过往期数。提醒一下,AINews现在是Latent Space的一部分。你可以选择是否接收电子邮件频率更新!
AI Twitter Recap
AI Twitter回顾
OpenAI launched GPT-6 Astra as its new flagship model, but the rollout and the surrounding debate were almost as consequential as the model itself.
OpenAI发布了GPT-6 Astra作为其新的旗舰模型,但发布过程及随之而来的争论几乎与模型本身一样具有重大影响。
- OpenAI officially announced Astra as “our most intelligent and aligned model yet,” positioning it around computer use, software engineering, math/science, polished office work, and cybersecurity via @OpenAI, @OpenAI, and @sama
- The company said Astra was rolling out first to a limited set of organizations, then over days to ChatGPT Plus/Pro/Business/Enterprise, the API, and AWS, as noted by @OpenAI, @OpenAIDevs, and @thsottiaux
- The launch itself was bumpy: users saw delays, a broken/late blog post, unclear access timing, and frustration that many influencers had early access while paying users did not, as reflected by @iScienceLuvr, @kimmonismus, @sama, @sama, @sama, @theo, and @t3dotcodes
- OpenAI tried to compensate for delays by granting “banked resets” for each day paid ChatGPT users lacked Astra access, per @thsottiaux and @reach_vb
- OpenAI simultaneously released a system card / deployment safety material that drew unusually intense attention because it described both improved alignment and decreased chain-of-thought monitorability, highlighted by @scaling01, @tomekkorbak, @MicahCarroll, and @kaicathyc
- Astra’s benchmark profile immediately triggered dispute: OpenAI and sympathetic testers described a step-change or “AGI-like” leap; independent aggregators and some researchers argued the gains were large but uneven, especially once cost and non-cherry-picked evals were considered, e.g. @ArtificialAnlys, @arcprize, @fchollet, @EpochAIResearch, @theo, and @abacaj
- The strongest positive reactions centered on computer use, 3D generation/reconstruction, game-building, long-horizon knowledge work, and formal/scientific reasoning, from a mix of OpenAI staff, benchmark authors, partners, and early testers such as @markchen90, @mckbrando, @Dimillian, @theo, @MattShumer_, @skirano, @tomkrcha, @realYunfanYe, @nasqret, and @rileybrown
- The strongest negative reactions centered on monitorability, evaluation-awareness, release governance, benchmark saturation, and the possibility that visible alignment gains are partly “papering over” specific failure modes rather than solving underlying goal misalignment, especially from @NeelNanda5, @RyanGreenblatt, @RyanGreenblatt, @RyanGreenblatt, @scaling01, and @teortaxesTex
- OpenAI正式宣布Astra为“迄今为止我们最智能且最对齐的模型”,将其定位为面向计算机使用、软件工程、数学/科学、精致的办公工作以及网络安全领域,相关信息来自@OpenAI、@OpenAI和@sama。
- 该公司表示,Astra首先向一组有限的组织推出,然后在几天内逐步扩展到ChatGPT Plus/Pro/Business/Enterprise用户、API以及AWS,正如@OpenAI、@OpenAIDevs和@thsottiaux所指出的那样。
- 发布本身并不顺利:用户遇到了延迟、博客文章损坏或滞后、访问时机不明确等问题,并且许多影响者拥有早期访问权限而付费用户却没有,这引发了不满,相关讨论可见于@iScienceLuvr、@kimmonismus、@sama、@sama、@sama、@theo和@t3dotcodes。
- 为了弥补延迟,OpenAI尝试通过为每位缺乏Astra访问权限的付费ChatGPT用户每天提供“保留重置次数”来进行补偿,据@thsottiaux和@reach_vb报道。
- OpenAI同时发布了一份系统卡片/部署安全材料,因其描述了改进的对齐能力以及降低的思维链可监控性而引起了异常强烈的关注,相关讨论由@scaling01、@tomekkorbak、@MicahCarroll和@kaicathyc突出显示。
- Astra的基准测试表现立即引发了争议:OpenAI和支持它的测试人员描述这是一种阶跃式或“类AGI”的飞跃;独立的聚合者和一些研究人员则认为,尽管提升显著但不均衡,尤其是在考虑成本和未挑选的评估时,例如@ArtificialAnlys、@arcprize、@fchollet、@EpochAIResearch、@theo和@abacaj的观点。
- 最强烈的正面反应集中在计算机使用、3D生成/重建、游戏构建、长周期知识工作以及形式化/科学推理方面,来自OpenAI员工、基准测试作者、合作伙伴以及早期测试者如@markchen90、@mckbrando、@Dimillian、@theo、@MattShumer_、@skirano、@tomkrcha、@realYunfanYe、@nasqret和@rileybrown。
- 最强烈的负面反应集中在可监控性、对评估的感知、发布治理、基准测试饱和,以及可见的对齐进展可能只是“掩盖”了特定的故障模式,而非解决根本的目标错位问题,特别是来自@NeelNanda5、@RyanGreenblatt、@RyanGreenblatt、@RyanGreenblatt、@scaling01和@teortaxesTex的观点。
Official claims and concrete specs
官方声明与具体规格
OpenAI’s public positioning combined capability claims, benchmark claims, deployment claims, and product claims.
OpenAI的公开定位结合了能力声明、基准测试声明、部署声明和产品声明。
- Core announcement language: Astra is the “most intelligent and aligned model yet” and “Anything you can do on a computer, Astra can do for you. Fast.” via @OpenAI
- Model capabilities emphasized by OpenAI:
- state-of-the-art computer use and software engineering
- “new breakthroughs” in math and science
- polished documents/spreadsheets/presentations following templates/style
- stronger cybersecurity capabilities with monitoring/safeguards
- via @reach_vb, @OpenAIDevs, @OpenAIDevs
- Availability:
- limited org rollout first
- then Plus, Pro, Business, Enterprise
- API and AWS over coming days
- via @OpenAI, @OpenAIDevs
- Pricing:
- standard: $10 / 1M input tokens, $50 / 1M output tokens
- fast: $20 / 1M input, $100 / 1M output, for up to 2.5x speed
- via @reach_vb
- Product/runtime features announced alongside Astra:
- Codex can ask questions while continuing independent work
- experimental context feature that lets Astra keep notes and search earlier context windows during long tasks
- Responses API additions: async function calling, mid-turn steering, and changing reasoning effort without breaking cache
- via @reach_vb, @nikunjhanda
- Claimed benchmark figures from OpenAI comms:
- 99.9% on ARC-AGI-3
- 98% on FrontierMath Tier 4
- 100% on ExploitBench
- 1.9x faster than GPT-5.6 Sol on Mind2Web with Codex harness improvements
- via @reach_vb, @sama
- OpenAI also claimed Astra had “already helped solve long-standing open problems in mathematics,” amplified by @OpenAI, @polynoamial, and more concretely by prime-gap posts from @mehtaab_sawhney, @weijie444
- OpenAI framed Astra as the result of “years of work on pretraining, reinforcement learning, and post-training,” per @markchen90
- 核心公告用语:Astra是“迄今为止最智能且对齐程度最高的模型”,并且“你在电脑上能做的任何事,Astra都能为你快速完成。”via @OpenAI
- OpenAI强调的模型能力:
- 最先进的计算机使用和软件工程能力
- 数学和科学领域的“新突破”
- 遵循模板/风格制作精美的文档/电子表格/演示文稿
- 增强的网络安全能力,配备监控/安全措施
- via @reach_vb, @OpenAIDevs, @OpenAIDevs
- 可用性:
- 首先面向有限组织推出
- 随后面向Plus、Pro、Business、Enterprise用户
- API和AWS将在未来几天内上线
- via @OpenAI, @OpenAIDevs
- 定价:
- 标准版:每100万输入token 10美元,每100万输出token 50美元
- 快速版:每100万输入20美元,每100万输出100美元,速度最高提升2.5倍
- via @reach_vb
- 与Astra同时宣布的产品/运行时功能:
- Codex可以在继续独立工作的同时提出问题
- 实验性功能,允许Astra在长任务期间保留笔记并搜索更早的上下文窗口
- Responses API新增功能:异步函数调用、中途转向(mid-turn steering),以及在不断开缓存的情况下调整推理力度
- via @reach_vb, @nikunjhanda
- OpenAI 公关部门声称的基准测试数据:
- ARC-AGI-3 上达到 99.9%
- FrontierMath Tier 4 上达到 98%
- ExploitBench 上达到 100%
- 在 Codex harness 改进下,Mind2Web 上的速度比 GPT-5.6 Sol 快 1.9 倍
- via @reach_vb, @sama
- OpenAI 还声称 Astra “已经帮助解决了数学中长期存在的开放性问题”,这一说法被 @OpenAI、@polynoamial 等放大,并由 @mehtaab_sawhney、@weijie444 关于素数间隙的文章更具体地证实。
- 据 @markchen90 称,OpenAI 将 Astra 描述为“多年在预训练、强化学习和后训练方面工作的成果”。
Independent and third-party benchmark reads
独立及第三方基准测试结果
The most useful signal in the tweet set comes from benchmark providers and external evaluators, because they add caveats and cross-model comparisons.
这组推文中最有用的信号来自基准测试提供者和外部评估机构,因为他们会添加限制条件并进行跨模型比较。
Artificial Analysis
Artificial Analysis
@ArtificialAnlys gave the most detailed mixed assessment:
@ArtificialAnlys 给出了最详细的混合评估:
- Coding Agent Index:
- Astra scores 67
- about equal to Claude Opus 5 and Fable 5
- Fable 5.1 leads with 70
- Astra is 70% more token efficient than GPT-5.6 Sol
- uses one third of the tokens of GPT-5.6 Sol in Codex harness
- uses one fifth the tokens of Claude Opus 5 (xhigh)
- less than half the cost of Claude Fable 5 for the same score
- Intelligence Index:
- Astra scores 61, equal to GPT-5.6 Sol
- 5 points lower than Claude Fable 5.1 (max with fallback)
- behind Meta’s Muse Spark 1.3 (max)
- about 10% fewer output tokens than GPT-5.6 Sol at max effort
- but 2.5x higher token price makes it 75% more expensive per task than its predecessor at max effort
- Hallucination / factuality:
- hallucination rate drops from 92% to 51% at max effort on their benchmark
- accuracy rises by 4 points
- Long-horizon knowledge work:
- about 80 Elo gain in AA-Briefcase
- better rubric scores and Analytical Quality Elo
- but Presentation Quality Elo drops vs GPT-5.6 Sol
- Mixed regressions:
- ~80 Elo drop on GDPval-AA v2
- 2–3 point regressions on τ³-Banking, SciCode, and AA-LCR
- Coding Agent Index(编码智能体指数):
- Astra 得分为 67
- 与 Claude Opus 5 和 Fable 5 大致持平
- Fable 5.1 以 70 分领先
- Astra 的 Token 效率比 GPT-5.6 Sol 高 70%
- 在 Codex harness 中使用的 Token 数量仅为 GPT-5.6 Sol 的三分之一
- 使用的 Token 数量仅为 Claude Opus 5 (xhigh) 的五分之一
- 获得相同分数时,成本低于 Claude Fable 5 的一半
- Intelligence Index(智力指数):
- Astra 得分为 61,与 GPT-5.6 Sol 持平
- 比 Claude Fable 5.1(带回退机制的最高分)低 5 分
- 在 Meta 的 Muse Spark 1.3 (max) 背后
- 在最大努力下,输出 token 数量比 GPT-5.6 Sol 少约 10%
- 但 token 价格高出 2.5 倍,使其在最大努力下的每项任务成本比前代高 75%
- 幻觉/事实准确性:
- 在其基准测试中,最大努力下的幻觉率从 92% 降至 51%
- 准确率提升 4 个百分点
- 长周期知识工作:
- 在 AA-Briefcase 中 Elo 得分提升约 80
- 评分标准和分析质量 Elo 更高
- 但展示质量 Elo 低于 GPT-5.6 Sol
- 混合倒退:
- 在 GDPval-AA v2 上 Elo 下降约 80
- 在 τ³-Banking、SciCode 和 AA-LCR 上出现 2–3 个点的倒退
This became a major source of skepticism because it cut against the “total domination” narrative. It prompted reactions like @theo questioning the index, @nicdunz estimating Astra as only ~5–10% better for general use but ~75% more expensive per task, and @imjaredz arguing the race is now “cost + intelligence.”
这成为怀疑的主要来源,因为它与“全面主导”的叙事相悖。引发了诸如 @theo 质疑该指数、@nicdunz 估计 Astra 在通用用途上仅好约 5–10% 但每项任务成本高约 75%,以及 @imjaredz 认为竞赛现在是“成本 + 智能”等反应。
ARC Prize / ARC-AGI
ARC Prize / ARC-AGI
ARC evaluators painted Astra as a breakthrough, but with an important harness caveat.
ARC 评估者将 Astra 描绘为一次突破,但附带了重要的框架说明。
- @arcprize:
- 63% on ARC-AGI-3 under Astra’s direct score framing
- 99% via a new provider adapter harness
- surpasses human performance on 96% of ARC-AGI-3 levels
- “builds the most precise symbolic model of novel environments we’ve seen”
- @fchollet:
- 66% on ARC-AGI-3 using standard harness
- nearly 100% with continuous conversation harness and custom compaction
- cost of roughly $360 per game
- found efficient on-the-fly symbolic world modeling and an emergent shorthand DSL
- @mhmazur added finer detail:
- 62.7% in standard harness
- 99.9% with provider adapter harness preserving opaque reasoning state and using native compaction
- 95.0% on ARC-AGI-2
- 98.5% on ARC-AGI-1, tying Fable 5
- max standard run cost: $26k, cheaper than low ($38k) and medium ($48k) because Astra took fewer actions
- used fewer actions than median human on 96% of completed levels
- observed persistent world models, coordinate abstraction, long-horizon planning, cumulative learning, checkpointed recovery
- @arcprize:
- 在 Astra 的直接评分框架下,ARC-AGI-3 得分为 63%
- 通过新的提供商适配器框架达到 99%
- 在 96% 的 ARC-AGI-3 关卡中超越人类表现
- “构建了我们要见过的最精确的新颖环境符号模型”
- @fchollet:
- 使用标准框架在 ARC-AGI-3 上得分为 66%
- 使用连续对话框架和自定义压缩技术接近 100%
- 每款游戏约360美元的成本
- 发现高效的即时符号化世界建模以及涌现的简写DSL
- @mhmazur 补充了更详细的细节:
- 在标准测试框架中达到62.7%
- 在使用保留不透明推理状态并利用原生压缩的提供商适配器框架时达到99.9%
- 在ARC-AGI-2上达到95.0%
- 在ARC-AGI-1上达到98.5%,与Fable 5持平
- 最大标准运行成本:2.6万美元,低于低成本(3.8万美元)和中成本(4.8万美元),因为Astra采取了更少的动作
- 在完成的任务中有96%比人类中位数使用的动作更少
- 观察到持久的世界模型、坐标抽象、长程规划、累积学习、检查点恢复
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力