跳到主内容
@wquguru
精选70Rohan Paul模型发布/更新多源精选 ×3

ARC-AGI-3基准饱和:GPT-6 Astra得分差异源于上下文管理方式

ARC-AGI-3 is now saturated, GPT-6 Astra reaches 99% with a new provider-adapter…

原文
发到 X

ARC-AGI-3 is now saturated, GPT-6 Astra reaches 99% with a new provider-adapter harness and beats human performance on 96% of ARC-AGI-3 levels.

ARC-AGI-3 现已饱和,GPT-6 Astra 借助新的 provider-adapter harness 达到 99%,并在 ARC-AGI-3 的 96% 的关卡中超越人类表现。

A benchmark designed to expose failures in novel reasoning is now close to a ceiling for the strongest system.

一项旨在揭示新型推理失败情况的基准测试,目前最强大的系统已接近天花板。

But the interesting part is that, the exact same model looks radically different depending on how the evaluation system manages its memory between actions.

但有趣的部分在于,完全相同的模型,根据评估系统在动作之间如何管理其记忆,看起来会有根本性的差异。

With ARC Prize's provider-neutral Standard harness, GPT-6 Astra tops out at 62.7%. Give Astra OpenAI's native way of preserving reasoning state and managing long context, and it jumps to 99.9%. The model did not suddenly get new weights. The environment did not become easier. The evaluation wrapper changed.

使用 ARC Prize 的与提供商无关的标准 harness(Standard harness),GPT-6 Astra 最高仅达 62.7%。若让 Astra 使用 OpenAI 原生保留推理状态和管理长上下文的机制,其得分则跃升至 99.9%。该模型并未突然获得新权重,环境也未变得更容易。变化的是评估包装器(evaluation wrapper)。

So "Provider Adapter harness" basically means:

因此,“Provider Adapter harness”基本上意味着:

Run the ARC benchmark while letting the model use the context-management machinery its own provider designed for it.

在运行 ARC 基准测试时,允许模型使用其自身提供商为其设计上下文管理机制。

That is different from ARC's Standard harness, which deliberately tries to provide a minimal, common interface that works similarly across different model providers. ARC describes these as answering 2 different questions: the Standard harness asks how models compare under the same provider-neutral setup, while the Provider Adapter asks how well a model performs when used with the context features its provider built for it.

这与 ARC 的 Standard harness 不同,后者刻意试图提供一个最小化的通用接口,使其在不同模型提供商之间以相似的方式工作。 ARC 将这两者描述为回答两个不同的问题:Standard harness 询问在相同的与提供商无关的设置下模型之间的比较情况,而 Provider Adapter 则询问当模型与其提供商为其构建的上下文功能配合使用时,其表现如何。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →