跳到主内容
精选80Rohan Paul模型发布/更新多源精选 ×4

Grok 4.6发布,Agent能力领先,编码略逊

Grok 4.6 just dropped.

原文

Grok 4.6 just dropped.

Clearest strength is professional agent work, with leading results against GPT Sol Max and Fable 5 Max on GDPVal-AA v2 and AA-Briefcase.

  • matches GPT-5.6 Sol Max at 61 on Artificial Analysis while charging $2/$6 per million input/output tokens.
  • 1753 on GDPVal-AA v2 and 1577 on AA-Briefcase, both ahead of Sol Max and Fable 5 Max. Those benchmarks measure real-world agent tasks and agentic knowledge work, making them closer to research, analysis and multi-file deliverables than isolated question answering.

On coding, 69.9% on CursorBench v3.2 beats Sol's 67.2%, while DeepSWE and Terminal-Bench leave Grok behind both Sol and Fable.

SpaceXAI attributes the jump to a longer training run, regenerated SFT trajectories, model-based trace filtering, and agentic RL across coding, web development, CAD and kernel optimization.

It also reports more self-testing on long trajectories, with the model checking its work before continuing, directly targeting error accumulation across multi-step agents.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近