双华为昇腾卡运行Qwen3.8-Flash:vLLM适配与工程复盘
Two 96 GB Ascend cards crun Qwen3.8-flash-next hardware notes, vLLM work, benchmarks, and what is next
昇腾生态稀缺的一手落地指南,详细拆解了vLLM在国产卡上的适配难点与性能调优路径,做国产算力迁移的同学值得收藏参考。
I have been building a somewhat unusual local inference machine around two Huawei Atlas 300I Duo cards. They are relatively inexpensive, passive, dual-accelerator PCIe cards with 96 GB of device memory apiece. They are also absolutely not drop-in CUDA replacements.
我一直在围绕两张华为 Atlas 300I Duo 卡构建一台相当特殊的本地推理机器。它们相对便宜,是无风扇的双加速器 PCIe 卡,每张拥有 96 GB 的设备内存。但它们绝对无法直接作为 CUDA 的替代品使用。
When I first brought up Qwen3.8 Flash-Next these past two weeks, it was often incoherent and lived around 1 generated token per second. Some runs were below that. Today the same two-card machine is producing coherent output at roughly 30 tok/s for one request and about 61 tok/s aggregate at four-way concurrency on my short decode benchmark. It also completed the full 198-question GPQA Diamond set.
在过去两周里,当我首次启动 Qwen3.8 Flash-Next 时,它经常输出混乱的内容,生成速度约为每秒 1 个 token(generated token per second)。有些运行甚至低于这个速度。今天,同样的双卡机器在我的短解码基准测试中,对于单个请求能产生连贯的输出,速度约为每秒 30 个 token;在四路并发下聚合速度约为每秒 61 个 token。它还完成了完整的 198 题 GPQA Diamond 数据集。
This post is the start of a guide for these cards: what the cards physically are, how I cool them, what “96 GB” really means, what I changed in vLLM and vLLM Ascend, which optimizations actually mattered, and which problems are still open.
本文是这些卡片的指南开篇:介绍这些卡片的物理形态、我的散热方案、“96 GB”究竟意味着什么、我对 vLLM 和 vLLM Ascend 做了哪些修改、哪些优化真正起了作用,以及哪些问题仍未解决。
The short version is the hardware is capable. My work has been mostly on the software stack, ubuntu-26.04 driver support, model architecture support, memory layout, custom operators, and getting every asynchronous state transition exactly right.
简而言之,硬件本身是具备能力的。我的工作主要集中在软件栈上,包括 ubuntu-26.04 驱动支持、模型架构支持、内存布局、自定义算子,以及确保每个异步状态转换完全正确。
The hardware: one card is really two devices
硬件方面:一张卡实际上是两个设备
My current machine has two Atlas 300I Duo cards, which enumerate as four Ascend 310P3 devices.
我当前的机器有两张 Atlas 300I Duo 卡,它们被枚举为四个 Ascend 310P3 设备。
Current system / Planned system
当前系统 / 计划系统
Physical cards: 2 / 3
物理卡片:2 / 3
Ascend devices/npus/AIcpus: 4 / 6
Ascend 设备/npus/AIcpus:4 / 6
Nameplate device memory: 192 GB / 288 GB
铭牌设备内存:192 GB / 288 GB
Approx. runtime-visible memory with this configuration: 172 GiB / 258 GiB
此配置下运行时可见的大致内存:172 GiB / 258 GiB
Combined maximum accelerator-board power: 300 W / 450 W
组合最大加速板功耗:300 W / 450 W
Each card has two accelerator SoCs and 96 GB of LPDDR4X in total, or 48 GB local to each chip. It is not one unified 96 GB allocation. A model that does not fit on one 48 GB device still needs tensor, expert, pipeline, or another form of model parallelism. The card is PCIe Gen4 x16, full-height/full-length, and a surprisingly thin single-slot design. Huawei rates it at 408 GB/s aggregate memory bandwidth and 150 W maximum board power. The official specifications are here (https://support.huawei.com/enterprise/en/doc/EDOC1100285916?section=j00e).
每张卡有两个加速器 SoC 和总共 96 GB 的 LPDDR4X 内存,或者每个芯片本地拥有 48 GB。这并不是一个统一的 96 GB 分配。如果一个模型无法装入单个 48 GB 的设备,仍然需要张量并行、专家并行、流水线并行或其他形式的模型并行。该卡采用 PCIe Gen4 x16 接口,全高全长设计,且出乎意料地采用了单槽薄型设计。华为标称其聚合内存带宽为 408 GB/s,最大板载功耗为 150 W。官方规格见此处 (https://support.huawei.com/enterprise/en/doc/EDOC1100285916?section=j00e)。
Also, despite the generic “HBM” terminology used by a lot of accelerator software, the memory on these cards is LPDDR4X.
此外,尽管许多加速器软件使用了通用的“HBM”术语,但这些卡上的内存实际上是 LPDDR4X。
This two-chip-per-card layout matters. Communication within a model still goes through the distributed runtime, and memory remains local to a rank. I use HCCL collectives and explicitly map tensor and expert ownership across all four chips. Thinking of the machine as four 48 GB ranks is much more useful than thinking of it as two 96 GB GPUs.
这种每张卡双芯片的布局很重要。模型内部的通信仍然通过分布式运行时进行,内存也保持在各 rank 本地。我使用 HCCL 集合操作,并明确映射所有四个芯片上的张量和专家所有权。将机器视为四个 48 GB 的 rank,比将其视为两个 96 GB 的 GPU 要更有用得多。
https://reddit.com/link/1wvt1m4/video/7rju7qbrw1th1/player
Passive cooling is not a deal-breaker
被动散热并不是致命缺陷
The cards have large heatsinks and no onboard fans. They were designed for server airflow, so putting them in an ordinary workstation and hoping a rear case fan will sort it out is a bad plan. There is a useful teardown here (https://videocardz.com/newz/huawei-atlas-300i-dual-ai-gpu-with-96gb-memory-worth-1400-has-been-taken-apart) if you want to see the heatsink and heat-pipe arrangement.
这些显卡配备了大型散热器且没有板载风扇。它们是为服务器气流设计的,因此将它们放入普通工作站并指望后置机箱风扇能解决问题是个糟糕的计划。如果你想查看散热器和热管布局,这里有一个有用的拆解视频(https://videocardz.com/newz/huawei-atlas-300i-dual-ai-gpu-with-96gb-memory-worth-1400-has-been-taken-apart)。
I give them direct, high-volume airflow and run the room on AC/heat-pump cooling. Under real multi-hour model loads, the cards can crank continuously without drama. Across my recorded Qwen runs, peak device temperatures were generally 72–78 °C. My watchdog limit is 96 °C, and the cards have not approached it.
我为它们提供直接、大流量的气流,并使用空调/热泵冷却房间。在实际的多小时模型负载下,这些显卡可以持续满负荷运行而不会出现任何问题。在我记录的 Qwen 运行中,设备峰值温度通常在 72–78 °C 之间。我的看门狗限制是 96 °C,而这些显卡从未接近过该阈值。
So I would not bat an eye at adding another passive card. The actual checklist is mundane:
因此,我不会对添加另一张被动散热显卡感到惊讶。实际的检查清单很平常:
• Keep unobstructed airflow through the heatsink fins.
• 确保气流不受阻碍地穿过散热器鳍片。
• Make sure the chassis fans have enough static pressure.
• 确保机箱风扇具有足够的静压。
• Budget another 150 W of board power per card, plus the rest of the host.
• 为每张卡预留另外 150 W 的板卡功耗,再加上主机其余部分的功耗。
• Exhaust the heat from the room instead of recirculating it through the rack.
• 将热量排出室外,而不是通过机架循环再吸入。
• Log temperature during long prefill, decode, and concurrency tests rather than trusting an idle reading.
• 在长时间的预填充(prefill)、解码(decode)和并发测试期间记录温度,而不是依赖空闲读数。
Passive does not mean low-power or self-cooling. It means the chassis and room are the cooling system. Once that is handled, these have behaved like ordinary 150 W server cards for me.
被动散热并不意味着低功耗或自冷。它意味着机箱和房间构成了冷却系统。一旦处理好这一点,这些显卡对我来说就像普通的 150 W 服务器显卡一样表现良好。
Note: Nothing heats up these cards more than loading/moving things around in their ram -- the npus at full utilization run cooler than large block memory assignments. We keep this in mind when optimizing the model serving code paths.
注意:没有什么比在显存中加载/移动数据更让这些显卡升温的了——NPU 在全利用率下运行时的温度反而低于大块内存分配时。我们在优化模型服务代码路径时会牢记这一点。
ECC, nameplate memory, and what is actually usable
ECC、标称内存以及实际可用的内存
My cards arrived with ECC enabled by default. I disabled it to reclaim device memory. This is an inference and development box, and I consciously prefer capacity over ECC protection here. That is a reliability tradeoff, not a universal recommendation.
我的显卡默认启用了 ECC。我禁用了它以回收设备内存。这是一台推理和开发用的机器,我在这里有意识地优先选择容量而非 ECC 保护。这是一种可靠性权衡,并非普遍推荐。
Even with ECC disabled, firmware, the runtime, communication buffers, graph captures, workspaces, and allocator reservations consume memory. In practice, the software sees roughly 43 GiB per 48 GB chip. Four chips therefore provide about 172 GiB of useful aggregate capacity, but it is still four separate local pools. The exact free number also changes with the CANN build and launch configuration.
即使禁用了 ECC,固件、运行时、通信缓冲区、图捕获、工作区和分配器预留仍会消耗内存。实际上,软件在每个 48 GB 芯片上可见的容量约为 43 GiB。因此,四个芯片总共提供约 172 GiB 的有效聚合容量,但它们仍然是四个独立的本地池。确切可用数量也会随 CANN 构建和启动配置而变化。
That distinction has shaped nearly every model decision. The question is not only, “Does the checkpoint total fit in 192 GB?” It is, “Does each rank's weight shard, recurrent state, KV/cache allocation, graph capture, collective workspace, and worst-case temporary allocation fit in its own 43 GiB?”
这一区别塑造了几乎所有的模型决策。问题不仅仅是“检查点总量是否适合 192 GB?”而是“每个 rank 的权重分片、循环状态、KV/缓存分配、图捕获、集体工作区以及最坏情况下的临时分配是否适合其各自的 43 GiB?”
Why I forked vLLM as well as vLLM Ascend
我为何分叉 vLLM 以及 vLLM Ascend
The public work lives in the OpenSensor vLLM Ascend fork (https://github.com/opensensor/vllm-ascend) and paired vLLM fork (https://github.com/opensensor/vllm). I needed both sides because this was not just a missing device kernel.
公开的工作位于 OpenSensor vLLM Ascend 分叉(https://github.com/opensensor/vllm-ascend)和配套的 vLLM 分叉(https://github.com/opensensor/vllm)。我需要这两部分,因为这不仅仅是一个缺失的设备内核。
I am currently the only person developing these forks. The software bus factor today is one. I have made a lot of progress, but a fast-moving one-person fork should not be confused with the maturity, test coverage, or support depth of mainline vLLM on NVIDIA.
我目前是唯一在开发这些分叉的人。当前的软件关键人物系数(bus factor)为 1。我已经取得了大量进展,但一个快速迭代的单人分叉不应与 NVIDIA 主线 vLLM 的成熟度、测试覆盖率或支持深度相混淆。
I also ran into a bizarre tooling problem: in my sessions, Claude repeatedly refused to engage with prompts about this architecture because the cards are Huawei hardware from China. These were ordinary engineering discussions about serving, sharding, cooling, and performance—not requests to build a restricted application. I am describing my direct experience rather than claiming that every Claude version or account will behave identically, but it made Claude unreliable as a development assistant for this project. Whatever anyone thinks about the politics, that is a real practical constraint when choosing tools around this hardware.
我还遇到了一个奇怪的工具问题:在我的会话中,Claude 反复拒绝处理关于此架构的提示,因为该硬件是中国华为的产品。这些都是关于服务、分片、冷却和性能的普通工程讨论——并非请求构建受限应用程序。我是在描述我的直接经验,而非声称每个 Claude 版本或账户都会表现一致,但这使得 Claude 作为该项目开发助手变得不可靠。无论人们如何看待政治因素,在选择围绕此硬件的工具时,这确实是一个现实的实际约束。
Qwen3.8 Flash-Next combines MoE routing, Gated DeltaNet recurrent layers, sparse quadratic-attention layers, packed low-bit experts, long context, and an MTP draft model. Supporting that cleanly touched model integration, the v1 runner, cache accounting, scheduling, graph capture, distributed state, model loading, and Ascend-specific operators.
Qwen3.8 Flash-Next 结合了 MoE 路由、门控 DeltaNet 循环层、稀疏二次注意力层、打包的低比特专家、长上下文以及 MTP 草稿模型。干净地支持这一点涉及模型集成、v1 运行器、缓存记账、调度、图捕获、分布式状态、模型加载以及 Ascend 特定算子。
The current development sprint has been roughly two and a half weeks of nearly continuous bring-up and optimization, with hundreds of fork commits, repeated full checkpoint loads, profiler captures, operator microbenchmarks, and multi-hour quality runs. This was not one magic kernel patch.
当前的开发冲刺经历了大约两周半近乎不间断的启动与优化工作,期间包含数百次 fork 提交、重复的全量检查点加载、性能分析器捕获、算子微基准测试以及持续数小时的质量运行。这并非依靠某个神奇的内核补丁就能实现。
What took Qwen from incoherent ~1 tok/s to where it is now
是什么让 Qwen 从混乱的约 1 tok/s 提升到现在的水平?
These are the architectural changes that moved the needle.
正是这些架构上的改变带来了显著进展。
- Make the hybrid model correct before making it fast
- 在追求速度之前,先确保混合模型的准确性
The early model could generate tokens, but generation was not a correctness test. I found failures that only appeared at production geometry: incorrect Gated DeltaNet gate-vector handling, recurrent-state precision and lifecycle problems, incomplete sparse-attention score width, and mismatches between the host operator API and the installed kernel package.
早期的模型能够生成 token,但生成过程本身并不能作为正确性测试的依据。我在生产环境几何结构下发现了仅在此场景才会出现的故障:包括 Gated DeltaNet 的门控向量处理错误、循环状态精度与生命周期问题、稀疏注意力分数宽度不完整,以及主机算子 API 与已安装的内核包之间的不匹配。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力