中国10T参数MoE模型传闻:30K Blackwell估算可信
~30K Blackwells sounds like a credible estimate for the rumored 10T-total-param…
~30K Blackwells sounds like a credible estimate for the rumored 10T-total-param MoE model from China.
约3万块Blackwell芯片听起来是对传闻中来自中国的10万亿总参数MoE模型的可信估计。
but parameter count isn't what sets it.
但参数数量并不是关键所在。
MoE training cost (FLOPs) ≈ 6 × active params × tokens.
MoE训练成本(FLOPs)≈ 6 × 激活参数 × 令牌数。
10T total is a memory bill, not a compute bill.
10万亿总参数是内存账单,而非计算账单。
Frontier sparsity is also collapsing:
前沿稀疏性也在崩溃:
DeepSeek-V3 671B/37B = 5.5% -> V4-Pro 1.6T/49B = 3.1% -> Kimi K3 routes 16 of 896 experts.
DeepSeek-V3 671B/37B = 5.5% -> V4-Pro 1.6T/49B = 3.1% -> Kimi K3 路由16/896个专家。
At 10T that's ~200–500B active.
在10万亿规模下,激活参数约为2000亿至5000亿。
Over 15–40T tokens -> 2e25 – 1.2e26 FLOP.
在15万亿至40万亿令牌上 -> 2e25至1.2e26 FLOP。
ByteDance already rents 36,000 Blackwell GPUs — 500 GB200 racks, $2.5 B. (through Malaysia cloud operator , deal is in line with US export controls).
字节跳动已租用36,000块Blackwell GPU——500个GB200机架,耗资25亿美元(通过马来西亚云运营商,交易符合美国出口管制)。
And Labs size clusters to the run they're planning.
实验室根据计划中的运行规模来调整集群大小。
And FT reports ByteDance’s pretraining window at 3-6 months.
金融时报报道字节跳动的预训练窗口为3至6个月。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力