跳到主内容
@wquguru
精选70Rohan Paul产品发布/更新

多智能体编排层性能超越Claude Code和Codex

So much recent work and research papers points to the same thing: the "harness"…

原文
发到 X

So much recent work and research papers points to the same thing: the "harness" is becoming the real capability layer.

@Offloop 's 4-person team demonstrated a multi-agent harness outperforming Claude Code and Codex on GDPval benchmarks, targeting $2.4T in US knowledge work.

  • The team scored 84.9 at $1.65 per task. - Opus 4.8 inside Claude Code scored 82.4, and GPT 5.6 Sol inside Codex scored 83.3, costing far more per task, $14.38 and $5.20

GDPval measures how well AI handles real work across 44 occupations and 9 major industries. Models get shell access and web browsing, then face blind pairwise comparisons against human experts.

And those tasks map onto US jobs paying roughly $2.4T a year.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近