跳到主内容
@wquguru
精选70François Chollet行业动态

François Chollet:静态基准测试本质衡量记忆而非智能

If your benchmark relies on a static dataset or sampling from a static distribut…

原文
发到 X

If your benchmark relies on a static dataset or sampling from a static distribution densely known at training time, then it is fundamentally measuring memorization/retrieval. Which might be fine if you're looking for a retrieval benchmark! But don't confuse it with intelligence.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
Grok 4.20 超越 Opus 4.8,Kimi K2.5 超 GLM 5.2,基准测试引争议
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文

相似阅读

另一事件,读法相近