跳到主内容
精选70Chubby♨️模型发布/更新多源精选 ×5

GLM-5.2 在真实工作基准 GDPval-AA 上排名第三,超越 GPT-5.5

Absolutely incredible: GLM-5.2 (max) sits at #3 overall on GDPval-AA, a real-wor…

原文

Absolutely incredible: GLM-5.2 (max) sits at #3 overall on GDPval-AA, a real-world agentic work benchmark, even ahead of GPT-5.5 (xhigh).

Oh and btw: looks like open source is no longer 7 months behind.

GDPval-AA, a benchmark built around real professional and creative tasks. The models had to produce practical deliverables from identical briefs, including a retail supervisor’s task list, an emergency-stop circuit schematic, and a music video moodboard.

Thats why we'll probably see a big leap with GPT-5.6. Even open source competition is catching up insanley fast.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近