跳到主内容
@wquguru
精选88Epoch AI(RSS)行业动态

Epoch AI:华为AI芯片性能与产能落后英伟达约四年

How far behind Nvidia is Huawei?

原文
发到 X
推荐理由

深度量化了华为与英伟达在算力硬件上的真实差距,数据详实且逻辑严密,对判断国产算力供应链现状极具参考价值。

This is a summary of a more detailed report, available on our website.

这是对一份更详细报告的摘要,该报告可在我们的网站上查阅。

The US-China AI competition spans everything from models and data centers to chips and the manufacturing equipment used to make them. Chips and manufacturing equipment form the foundation of that stack and determine how much compute each side can field.

中美人工智能竞争涵盖从模型、数据中心到芯片及制造这些芯片所需设备的一切领域。芯片和制造设备构成了这一技术栈的基础,并决定双方能部署多少算力。

This is why semiconductor export controls have been a centerpiece of US AI policy, and why so much depends on how China’s homegrown chip industry stacks up against America’s.

这就是为什么半导体出口管制成为美国人工智能政策的核心,以及为什么局势在很大程度上取决于中国本土芯片产业与美国相比的表现。

Huawei, China’s leading AI chip designer, laid out its chip roadmap last year and showcased an accelerated version last week at its annual conference. The roadmap is ambitious: it plans to pack more compute into each chip, connect far more chips into a single high-speed system, and mature the software that determines how much of that hardware is actually used.

中国领先的AI芯片设计公司华为去年公布了其芯片路线图,并在上周的年会上展示了一个加速版本。该路线图雄心勃勃:计划在每个芯片中集成更多算力,将更多芯片连接成单个高速系统,并完善软件以确定实际利用了多少硬件。

But Huawei currently lags Nvidia on both per-chip performance and the number of chips produced. Multiply the two, and Huawei will produce roughly 25× less total compute than Nvidia in 2026. US export controls constrain its ability to improve either factor, while Nvidia continues to push the frontier.

但目前在单芯片性能和芯片产量方面,华为都落后于英伟达。将两者相乘,华为在2026年的总算力产出将约为英伟达的1/25。美国的出口管制限制了其在任一因素上提升的能力,而英伟达则继续推动技术前沿。

With Nvidia making strides, can Huawei close the gap?

随着英伟达取得进展,华为能否缩小差距?

After crunching the numbers, we think almost certainly not. Between now and 2030, Huawei will likely remain around four years behind in both chip performance and compute production.

经过计算,我们认为几乎肯定不能。从现在到2030年,华为在芯片性能和算力生产方面可能仍将落后约四年。

Huawei’s starting position in 2026

华为在2026年的起始位置

In 2026, Huawei is significantly behind Nvidia. To understand how far Huawei needs to go to catch up, we can compare Huawei and Nvidia on the total compute each produces, which is the number of chips each makes multiplied by the performance of those chips. We convert all performance measurements to the equivalent number of Nvidia H100 chips, called H100-equivalents (H100e).

在2026年,华为远远落后于英伟达。为了了解华为需要走多远才能赶上,我们可以比较华为和英伟达各自生产的总算力,即各自生产的芯片数量乘以这些芯片的性能。我们将所有性能测量值转换为等效的英伟达H100芯片数量,称为H100当量(H100e)。

Performance: Huawei’s flagship AI chip for 2026, the Ascend 950, delivers roughly 7× less compute throughput than Nvidia’s B300 and half that of Nvidia’s 2022 H100. In other words, Huawei’s latest and most powerful GPU delivers only half the performance of a two-generation-old Nvidia chip that first started shipping in 2022, suggesting that Huawei’s chip performance currently trails Nvidia’s by four years.1

性能:华为2026年的旗舰AI芯片Ascend 950提供的计算吞吐量比英伟达B300少约7倍,仅为英伟达2022年H100的一半。换句话说,华为最新且最强大的GPU仅能提供两代前英伟达芯片(最早于2022年开始出货)一半的性能,这表明华为目前的芯片性能落后英伟达四年。1

Volume: For 2026, we estimate that Huawei will produce 1.5 million chips — one-fourth of Nvidia’s roughly 6 million.

产量:对于2026年,我们估计华为将生产150万颗芯片——仅为英伟达约600万颗的四分之一。

Total compute: With one-fourth the chips and each chip one-seventh the performance, Huawei will produce roughly 25× less compute than Nvidia in 2026.

总算力:由于芯片数量只有英伟达的四分之一,且每颗芯片的性能只有英伟达的七分之一,华为在2026年的总算力产出将约为英伟达的1/25。

Huawei’s roadmap to improve performance by 2030

华为到2030年提升性能的路线图

Closing that gap would require major progress in both individual chips and the systems connecting them — and Huawei has laid out an ambitious roadmap to improve both.

缩小这一差距需要在单个芯片和连接它们的系统方面取得重大进展——华为已经制定了一份雄心勃勃的路线图,旨在同时提升这两方面的性能。

Chip performance

芯片性能

Huawei starts out at a disadvantage on chip performance because Huawei’s fabrication partner SMIC can’t pack transistors as densely as Nvidia’s partner TSMC can. Transistors are the building blocks of a chip’s performance, so greater transistor density allows engineers to pack more compute onto one chip.

在芯片性能方面,华为处于劣势,因为其代工合作伙伴中芯国际(SMIC)无法像英伟达的合作伙伴台积电(TSMC)那样实现如此高的晶体管密度。晶体管是芯片性能的基础构建模块,因此更高的晶体管密度允许工程师在一个芯片上集成更多的计算能力。

US export controls restrict sales to China of the advanced lithography machines needed to produce the densest chips, and China’s efforts to produce these machines domestically likely won’t materialize until after 2030.

美国出口管制限制了向中国出售生产最密集芯片所需的最先进光刻机,而中国自主生产这些设备的努力可能在2030年之后才能见效。

Huawei’s roadmap targets a 7x performance improvement over two Ascend generations. Since the Ascend chips can’t scale transistor density quickly enough, it must gain performance improvements from other methods.

华为的路线图目标是在两代昇腾(Ascend)芯片上实现7倍的性能提升。由于昇腾芯片无法足够快地提升晶体管密度,它必须通过其他方法获得性能改进。

Bigger chips: First, Huawei can make its chips more powerful by making them bigger and packing each with more or larger logic dies (the parts that perform the computation), but the trade-off is higher costs and more defects. Meanwhile, Nvidia is also expanding its packages, possibly with fewer defects given its more mature process.

更大的芯片:首先,华为可以通过增大芯片尺寸并在每个芯片中集成更多或更大的逻辑裸片(执行计算的部分)来增强芯片性能,但代价是成本更高且缺陷率增加。与此同时,英伟达也在扩大其封装规模,鉴于其更成熟的工艺,可能具有更少的缺陷。

Chip Architecture: Huawei’s updated Ascend roadmap optimizes performance for lower precision formats, likely to match what frontier workloads are moving towards. This architecture update will allow the Ascend chips to perform more calculations with the same amount of silicon, significantly improving the Ascend’s performance running frontier AI workloads.

芯片架构:华为更新的昇腾路线图针对低精度格式优化了性能,这可能与前沿工作负载的发展趋势相匹配。此次架构更新将使昇腾芯片在相同硅面积下执行更多计算,从而显著提升其在运行前沿AI工作负载时的性能。

Vertical stacking: Huawei’s big, long-term bet is to stack logic dies vertically to fit far more transistors within a given footprint than SMIC’s process can print in a single layer. Stacking also shortens communication time within the chip, which allows the chip to run faster with less power. However, this technique, which Huawei calls LogicFolding, won’t reach the Ascend line until 2030. Nvidia’s Feynman generation, expected in 2028, will reportedly stack at least two layers of its superior logic dies, two years before the first Ascend attempts to do so. This would compound the 2× advantage from transistor density into a 4× advantage.

垂直堆叠:华为的一项长期重大押注是将逻辑裸片垂直堆叠,以便在给定 footprint 内容纳比中芯国际单层工艺所能制造的晶体管多得多的数量。堆叠还能缩短芯片内部的通信时间,从而使芯片在更低功耗下以更高速度运行。然而,这项被称为“LogicFolding”的技术要到2030年才会应用于昇腾产品线。据报道,英伟达预计于2028年推出的Feynman世代将至少堆叠两层其更优越的逻辑裸片,比第一代昇腾尝试此技术早两年。这将把晶体管密度带来的2倍优势叠加为4倍优势。

Faster systems

更快的系统

Like all major chip designers, Huawei is focused on optimizing the entire computing system, rather than just the individual chip.

与所有主要芯片设计公司一样,华为专注于优化整个计算系统,而不仅仅是单个芯片。

A decade ago, frontier models could be trained on a handful of GPUs. Today, training frontier models requires tens of thousands of AI chips. At that scale, performance depends heavily on how quickly chips can communicate and how effectively software can coordinate work across them. Huawei is pursuing both levers.

十年前,前沿模型只需少量 GPU 即可训练。如今,训练前沿模型需要数万个 AI 芯片。在如此规模下,性能在很大程度上取决于芯片之间的通信速度以及软件协调跨芯片工作的效率。华为正同时在这两个杠杆上发力。

Larger systems: Huawei plans to connect more chips inside a single domain. A domain is a group of chips wired together tightly so that they behave more like one giant system, which reduces the time spent waiting on communication from other parts of the system.

更大规模的系统:华为计划在一个域(domain)内连接更多的芯片。域是一组紧密布线连接的芯片,它们表现得像一个巨大的系统,从而减少了等待系统其他部分通信的时间。

Nvidia’s domain focuses on having extremely fast links within one rack of 72 chips. Inside the rack, any chip can talk to any other chip at full speed. Since Nvidia’s superior chips can keep communication-intensive work within fewer chips, it doesn’t need to compromise on communication speed for a larger domain size.

英伟达的域侧重于在单个包含 72 个芯片的机架内实现极快的链路连接。在机架内部,任何芯片都能以全速与其他任何芯片通信。由于英伟达更优越的芯片能够将通信密集型工作限制在更少的芯片中,因此它无需为了扩大域的规模而在通信速度上做出妥协。

Huawei goes the other way, compensating for weaker chips with larger domains; its planned “SuperPoDs” will connect 8,192 chips. However, maintaining high-bandwidth simultaneous communication from thousands of chips comes with unfeasibly steep trade-offs. Instead, Huawei implements a communication hierarchy where closer chips have more bandwidth, and more distant chips have to go through several hops. Whether this will perform well depends on Huawei’s software getting better at keeping communication local. Unfortunately, it’s hard to say how well the SuperPoD will perform in practice because Huawei has not published key performance metrics.

华为则采取相反的策略,通过更大的域来弥补芯片性能的不足;其计划的“SuperPod”将连接 8,192 个芯片。然而,维持来自数千个芯片的高带宽并发通信会带来难以承受的权衡代价。取而代之的是,华为实施了一种通信层级结构,即距离较近的芯片拥有更高的带宽,而距离较远的芯片必须经过多次跳接。这种架构能否表现良好,取决于华为的软件能否更好地保持通信的局部性。不幸的是,很难说 SuperPod 在实际中表现如何,因为华为尚未公布关键的绩效指标。

Better software: The software that coordinates workloads across chips can move performance by more than 10×, and it’s currently one of Nvidia’s strongest levers. Huawei’s equivalent of Nvidia’s CUDA is called CANN, which was first released in 2018 and makes working with Ascend chips difficult. When DeepSeek tried to train its R2 model on Ascend chips, Huawei sent its own engineers in and still couldn’t complete a training run, prompting DeepSeek to revert to Nvidia. However, if Huawei can capture the feedback loop from having Chinese labs use its chips instead of Nvidia’s, the software gap could close fast. For now, though, that possibility is still nascent: CANN remains years behind CUDA, and while Chinese labs are increasingly using domestic chips, they still reach for Nvidia hardware when they can get it.

更好的软件:协调跨芯片负载的软件可以提升超过 10 倍的性能,目前这也是英伟达最强的杠杆之一。华为对标英伟达 CUDA 的产品称为 CANN,于 2018 年首次发布,但使用 Ascend 芯片仍然困难重重。当 DeepSeek 尝试在 Ascend 芯片上训练其 R2 模型时,华为派遣了自己的工程师介入,但仍无法完成一次训练运行,这促使 DeepSeek 回退到英伟达平台。然而,如果华为能够利用中国实验室转而使用其芯片而非英伟达芯片这一反馈循环,软件差距可能会迅速缩小。不过目前,这种可能性仍处于萌芽阶段:CANN 仍落后 CUDA 数年之久,尽管中国实验室越来越多地使用国产芯片,但在条件允许时,它们仍会首选英伟达硬件。

Huawei will remain three to four years behind in chip performance

华为在芯片性能上将落后三到四年

On paper, Nvidia’s B300 delivers roughly 7× more compute performance than Huawei’s Ascend 950 series. As Nvidia moves from its Blackwell chips to Rubin and Rubin Ultra, the gap between each company’s best chip will widen to roughly 9× by 2027. The first Ascend chip expected to come close to the B300 in peak performance is the 970, slated to ship in late 2028. Given that the B300 arrived in 2025, that would put Huawei roughly three to four years behind Nvidia at the chip level even out to 2029.

纸面上,英伟达的B300提供的计算性能约为华为昇腾950系列的7倍。随着英伟达从Blackwell芯片转向Rubin和Rubin Ultra,到2027年,两家公司最佳芯片之间的差距将扩大到约9倍。预计峰值性能最接近B300的首款昇腾芯片是970,计划于2028年底出货。鉴于B300于2025年问世,这意味着即使到2029年,华为在芯片层面仍落后英伟达三到四年。

Can Huawei make up the difference in volume?

华为能在出货量上弥补这一差距吗?

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件