GPT-6 Astra在MazeBench无代码赛道得分14%
Big score by GPT-6 Astra, 14% on MazeBench without Python, 7X Claude Fable 5.1's…
Big score by GPT-6 Astra, 14% on MazeBench without Python, 7X Claude Fable 5.1's 2%.
GPT-6 Astra 取得高分,在 MazeBench 上达到 14%(不含 Python),是 Claude Fable 5.1 的 2% 的 7 倍。
MazeBench specifically stresses long-horizon spatial reasoning, with puzzles that can require more than 100 moves and repeated interaction with a changing environment.
MazeBench 专门考验长程空间推理能力,其中的谜题可能需要超过 100 步操作,并需要与不断变化的环境进行反复交互。
So Astra seems able to hold a long plan in its head and execute it without relying on Python to solve the environment for it
因此,Astra 似乎能够在脑海中保持一个长期计划,并执行该计划,而不依赖 Python 来为其解决环境问题。
Its remaining failures on genuinely 3D puzzles show that the improvement is mainly in sustained planning, not complete spatial understanding.
它在真正 3D 谜题上的剩余失败表明,改进主要体现在持续规划能力上,而非完全的空间理解能力。
MazeBench hides 100 gems across more than 200 rooms, forcing agents to rotate views, manipulate boxes, recover from mistakes, and execute plans that can exceed 100 moves.
MazeBench 在超过 200 个房间中隐藏了 100 颗宝石,迫使智能体旋转视角、操作箱子、从错误中恢复,并执行可能超过 100 步的计划。
The no-Python track makes it so much harder because code-enabled agents can reverse-engineer the physics and run solver algorithms instead of carrying the full spatial plan themselves.
无 Python 赛道使得任务变得困难得多,因为启用代码的智能体可以逆向工程物理机制并运行求解算法,而不是自己承载完整的路径规划。
MazeBench's creator says Astra still surpassed GPT-5.6 Sol's 13% code-enabled run and often planned 10-20 moves ahead in batched actions.
MazeBench 的创建者表示,Astra 仍然超越了 GPT-5.6 Sol 使用代码实现的 13% 的成绩,并且经常在批量操作中提前规划 10-20 步。
That batching reportedly reduced a projected 3B-token trajectory to about 350M tokens, but the run still lasted more than 60 hours.
据报道,这种批处理方式将预计的 30 亿 token 轨迹减少到约 3.5 亿 token,但运行时间仍超过 60 小时。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力