精选60Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)模型发布/更新
GLM缩小编码评估差距,ARC等推理基准或同样脆弱
Hear me:
Hear me: People used to soyface about novel coding evals, where Chyna/open models were not just behind but garbage. GLM covered most of that gap. Now we look at combined metrics like ECI, or "pure reasoning" like ARC. I predict this, too, will prove to be surprisingly fragile.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力