跳到主内容
@wquguru
精选70Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)模型发布/更新

DeepSeek系列训练成本大幅下降,V4-Flash约66K GPU小时

DeepSeek-V1 used 300K H800 GPU-hours per 1T tokens

原文
发到 X

DeepSeek-V1 used 300K H800 GPU-hours per 1T tokens DeepSeek-V2 used 173K GPU-hours DeepSeek-V3 used 180K GPU-hours DeepSeek-V4-Flash… probably around 66K. For a total of ≈2M. Like V2-Coder. Their current best model is smaller and cheaper than the one from 26 months ago.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近