跳到主内容
@wquguru
精选70Rohan Paul行业动态

中国10T参数MoE模型传闻:30K Blackwell估算可信

~30K Blackwells sounds like a credible estimate for the rumored 10T-total-param…

原文
发到 X

~30K Blackwells sounds like a credible estimate for the rumored 10T-total-param MoE model from China.

约3万块Blackwell芯片听起来是对传闻中来自中国的10万亿总参数MoE模型的可信估计。

but parameter count isn't what sets it.

但参数数量并不是关键所在。

MoE training cost (FLOPs) ≈ 6 × active params × tokens.

MoE训练成本(FLOPs)≈ 6 × 激活参数 × 令牌数。

10T total is a memory bill, not a compute bill.

10万亿总参数是内存账单,而非计算账单。

Frontier sparsity is also collapsing:

前沿稀疏性也在崩溃:

DeepSeek-V3 671B/37B = 5.5% -> V4-Pro 1.6T/49B = 3.1% -> Kimi K3 routes 16 of 896 experts.

DeepSeek-V3 671B/37B = 5.5% -> V4-Pro 1.6T/49B = 3.1% -> Kimi K3 路由16/896个专家。

At 10T that's ~200–500B active.

在10万亿规模下,激活参数约为2000亿至5000亿。

Over 15–40T tokens -> 2e25 – 1.2e26 FLOP.

在15万亿至40万亿令牌上 -> 2e25至1.2e26 FLOP。

ByteDance already rents 36,000 Blackwell GPUs — 500 GB200 racks, $2.5 B. (through Malaysia cloud operator , deal is in line with US export controls).

字节跳动已租用36,000块Blackwell GPU——500个GB200机架,耗资25亿美元(通过马来西亚云运营商,交易符合美国出口管制)。

And Labs size clusters to the run they're planning.

实验室根据计划中的运行规模来调整集群大小。

And FT reports ByteDance’s pretraining window at 3-6 months.

金融时报报道字节跳动的预训练窗口为3至6个月。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近