跳到主内容
@wquguru
精选60Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)模型发布/更新

GLM缩小编码评估差距,ARC等推理基准或同样脆弱

Hear me:

原文
发到 X

Hear me: People used to soyface about novel coding evals, where Chyna/open models were not just behind but garbage. GLM covered most of that gap. Now we look at combined metrics like ECI, or "pure reasoning" like ARC. I predict this, too, will prove to be surprisingly fragile.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
GLM 5.2 开源模型逼近前沿,开源 AI 迎来新纪元
Two Minute Papers(YouTube)原文

相似阅读

另一事件,读法相近