GPT-6 Astra在ARC-AGI-3基准测试中展现推理能力跃迁
GPT-6 Astra represents a step-function change in model capability for interactiv…
第三方权威评估确认GPT-6 Astra在复杂推理任务上的突破性表现,揭示了模型内部符号建模的新行为模式,对理解下一代模型能力边界极具参考价值。
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
GPT-6 Astra 在交互式推理问题的模型能力上实现了阶跃式提升。使用我们的标准测试框架,它在 ARC-AGI-3 上得分 66%;若结合连续对话测试框架与自定义压缩技术,其得分接近 100%,每局游戏的成本约为 360 美元。
In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.
事实上,连续测试框架版本的性能在几乎所有关卡的动作效率上都显著优于我们的人类基线。当我们检查推理链以了解模型的运作方式时,发现它为每个游戏和关卡执行了高效、实时的符号化世界建模。它甚至发展出了一种自己的简写 DSL(领域特定语言)来表示游戏内情境——本质上是一种针对特定游戏的代数记法。
Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.
总体而言,Astra 展现出了符号化建模行为,这类行为此前仅见于复杂的测试框架中——因此,测试框架的能力正日益转移到模型本身之中。
We see Astra as a major breakthrough in model intelligence.
我们将 Astra 视为模型智能方面的一项重大突破。
Read our post on Astra and what these results mean: https://arcprize.org/blog/astra
阅读我们关于 Astra 及其结果意义的博文:https://arcprize.org/blog/astra
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力