精选70Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)模型发布/更新
DeepSeek系列训练成本大幅下降,V4-Flash约66K GPU小时
DeepSeek-V1 used 300K H800 GPU-hours per 1T tokens
DeepSeek-V1 used 300K H800 GPU-hours per 1T tokens DeepSeek-V2 used 173K GPU-hours DeepSeek-V3 used 180K GPU-hours DeepSeek-V4-Flash… probably around 66K. For a total of ≈2M. Like V2-Coder. Their current best model is smaller and cheaper than the one from 26 months ago.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力