DeepSeek 算力布局:异构集群与国产芯片路线
Wenfeng's investor call should be read skeptically. It has several curious contr…
Wenfeng's investor call should be read skeptically. It has several curious contradictions. He says he won't do chips "if possible" (has actually been working on chips for years). The part about the government not giving him a cent (got $150M investment). And… $3B tops for compute? At $25K/pop, that's 120K H200s, but *he can't take so much* – the sales cap is 75K/entity. Add 20K he has (and almost certainly those are heavily H800s, H20s, some H100s – garbage), it still isn't enough. And yet, his job listing "AI Computing Cluster Performance and Reliability Engineer" explicitly talks about a 100K cluster they need to keep operational, filled with "supernodes" among all else. "on your first day, you'll be managing a 100.000 card cluster". It is training-oriented; the listing is full of autism about whole-system stability (inference nodes can fail independently). There's a lot of concern about different "batches" of hardware, which also shouldn't be a big issue with H200s; it's a very well understood device from a mature production line. On the other hand, it says NCCL, NVLink, "Understanding of RDMA/InfiniBand/GPU architecture or performance tuning (a plus, but not necessary)", so I guess he's not lying that he intends to NVidia-maxx as much as possible (this also plays to DeepSeek's strengths, they have sublime mastery of Nvidia stack at this point, arguably more than Nvidia itself). And on the gripping hand, there are listings that mention Ascend and "CPU/GPU/NPU" ("High-performance operator/communication/compiler engineer", "Supercomputing Cluster R&D Engineer").
I think it's a tightly coupled heterogenous system that'll, for example, use Nvidia for weight updates and SuperPoDs for rollouts, in the way we can infer from V4 paper. As Zephyr notes, 950s are relatively less inefficient in inference, and the software is definitely way more mature https://x.com/zephyr_z9/status/2080368613064945803.
Eventually Ascends will be supplemented or supplanted with in-house Whale Compute which is even more inference-optimal, and as the availability of Huawei hardware grows and Nvidia's dries up, Ascends will probably get shifted to training as well.
In short, it's tight, but I think he may well have the equivalent of "200K 950s" by the end of 2026 Q4, and his people will be able to conduct serious research that leads to next-generation Whales. An OOM scaled-up pretraining will have to come later.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力