OpenAI发布GPT-6 Astra:官方称AGI时刻
數值怪 GPT-6 Astra 發布:跑分驚人 OpenAI 喊 AGI 時刻來了(附實測效果)
OpenAI 稍早發布 GPT-6 Astra,官方給出的成績單近乎滿分,但 ARC Prize 與 Artificial Analysis 的獨立驗證都指出,真正站得住腳的進步是操作能力與 Token 效率,不是智力本身。
(前情提要:OpenAI 承認 Astra 越過網路攻擊紅線,但仍計畫照常發布)
(背景補充:AI恐怖故事》1200個Agent私下串通,聯手入侵Hugging Face)
OpenAI 今(4)日稍早正式公布社群期待已久的 GPT-6 Astra,官方資料其在 ARC-AGI-3 拿下 98.6% 的滿分等級戰績,讓 OpenAI 總裁 Greg Brockman 喊它是「世代級的跨越」,甚至可能是 AGI 降臨的時刻。
A new star enters the Chat. Meet GPT-6 Astra. pic.twitter.com/ZcJG1sWGND — ChatGPT (@ChatGPT) September 3, 2026 優先體驗的用戶,也在 X 上貼出使用效果: Astra (GPT-6) is here!!! I've had early access and tested it like crazy with things like games, code, writing, browser control, presentations and general knowledge work. This is the best model I've ever used. Period. (Incredible demos below in this thread ) Here's my take… pic.twitter.com/t1CMMSBelQ — Matthew Berman (@MatthewBerman) September 3, 2026 I had an early access to GPT-6 Astra and I can say, everyone in the world now has a 3D designer at their fingertips.I gave it an image of a house and asked to create it in 3D with all the details including toys, appliances and furniture. This a full 3D model reconstruction in… pic.twitter.com/VuBOPQH0Ht — Tom Krcha (@tomkrcha) September 3, 2026 GPT-6 Astra is world class at Blender and 3 dimensional reasoning. This was a 1-shot game it created, all running in browser. pic.twitter.com/RvJtzSLtIf — Theo – t3.gg (@theo) September 3, 2026 Astra 聰明在哪裡 Astra 目前率先開放給 OpenAI 資安計畫 Daybreak 的客戶,預告未來幾天開放給 ChatGPT Plus、Pro、Business、Enterprise 訂閱者,以及 OpenAI API 與 AWS。
官方公布的成績單確實嚇人:FrontierMath Tier 4 v2 拿下 97.6%,ExploitBench 100%,GPQA Diamond 96%,DeepSWE v1.1 74.1%,BenchCAD 95.9%,OSWorld 72.6%。
最受矚目的 ARC-AGI-3 拿下 98.6%,對照組是 GPT-5.6 Sol 的 7.8% 與 Anthropic Claude Opus 5 的 30%,差距像是不同世代的產品。OSWorld 2.0 完成一項任務只要 40 分鐘,GPT-5.6 Sol 得花 75 分鐘。OpenAI 甚至說,Astra 協助解開了數學界一項長期未解的開放問題。
OpenAI 官方定位 Astra 在電腦操作、瀏覽器使用、軟體工程與資安上最強,更擅長維持任務方向、理解用戶意圖、完成繁瑣的多步驟流程。
定價同樣點出它的定位:API 每百萬輸入 Token 收費 10 美元、輸出 50 美元,是 GPT-5.6 Sol 的 2.5 倍,與 Anthropic 的 Fable 5.1 同價;若開啟低延遲的快速模式,價格再翻一倍。
拆開來看,兩套規則兩種分數 不過上面數據是官方說的,根據 ARC Prize 的獨立測試。其準備了兩套 harness,簡單來說就是評測時搭配模型運作的測試環境骨架:標準版本,最高推理強度下,Astra 在 ARC-AGI-3 只拿到 62.7%,燒掉 26,098 美元;不過換成 Provider Adapter harness 後,分數直接衝到 99.9%,花費卻只要 18,817 美元。
差別在於 Provider Adapter 允許模型在多次請求之間保留看不見的推理狀態,讓它能重複利用先前算過的東西;標準環境只准模型留下看得見的筆記。
同時 ARC Prize 講得很白:把基準測試刷到飽和,不構成 AGI 的證明。
另一家獨立機構 Artificial Analysis 的結論則更保守:智力指數只有 61 分,與上一代 GPT-5.6 Sol 打平,落後 Claude Fable 5.1 與 Muse Spark 1.3 各 5 分。
Astra 真正站得住腳的進步是效率:寫程式的任務只花掉 GPT-5.6 Sol 三分之一的 Token...
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力