跳到主内容
@wquguru
精选88SemiAnalysis(RSS)行业动态

SemiAnalysis深度解析GLM-5.3稀疏注意力与推理成本优化

How GLM5.3 Sparse Attention Affects HBM Memory Usage

原文
发到 X
推荐理由

这是一份极具价值的推理基础设施深度评测,不仅量化了不同芯片在特定负载下的真实成本优势,还详细拆解了稀疏注意力与内存管理的工程取舍,做Agent或大模型部署的同学值得收藏参考。

How Sparse Attention Affects DRAM/NAND Memory

稀疏注意力如何影响 DRAM/NAND 内存

How does sparse attention affect the TAM of memory, including HBM and NAND? Sparse attention selects top-k most relevant tokens to attend to, reducing the memory consumption and bandwidth requirements during the core Scaled Dot-Production Attention (SDPA) operation. However, the efficiency improvement doesn’t directly translate to overall memory savings in practice. Concretely, the top-k selection operation typically requires the full context to be in HBM, so sparse attention doesn’t eliminate the memory capacity bottleneck.

稀疏注意力如何影响包括 HBM 和 NAND 在内的内存 TAM?稀疏注意力会选择最相关的 top-k 个 token 进行关注,从而减少核心缩放点积注意力(SDPA)操作期间的内存消耗和带宽需求。然而,效率的提升并不能直接转化为实际中的整体内存节省。具体而言,top-k 选择操作通常要求完整的上下文都位于 HBM 中,因此稀疏注意力并不能消除内存容量瓶颈。

Sparse attention throughput is bottlenecked by memory capacity. Source: HiSparse

稀疏注意力的吞吐量受限于内存容量。来源:HiSparse

To overcome this limitation, the SGLang team designed HiSparse, a hierarchical memory system that proactively offloads KV cache entries from device HBM to host DRAM. HiSparse behaves like an LRU (Least Recently Used) cache, where it loads tokens from DRAM to HBM upon top-k selection cache miss, and it offloads tokens from HBM to DRAM based on the LRU eviction policy. To reduce the cache miss latency, HiSparse overlaps KV cache loading of layer N with the execution of layer N-1 (layer-wise overlapping), introduced in HiSparse’s prior work HiCache.

为了克服这一限制,SGLang 团队设计了 HiSparse,这是一种分层内存系统,能够主动将 KV cache 条目从设备 HBM 卸载到主机 DRAM。HiSparse 的行为类似于 LRU(最近最少使用)缓存,当发生 top-k 选择缓存未命中时,它会将 token 从 DRAM 加载到 HBM,并基于 LRU 驱逐策略将 token 从 HBM 卸载到 DRAM。为了降低缓存未命中的延迟,HiSparse 将第 N 层的 KV cache 加载与第 N-1 层的执行重叠(逐层重叠),这一特性在 HiSparse 的先前工作 HiCache 中引入。

Layer-wise overlapping. Source: CachedAttention

逐层重叠。来源:CachedAttention

With HiSparse, SGLang greatly boosts throughput at high concurrency and long context scenarios, at the cost of top-k cache miss I/O overhead.

借助 HiSparse,SGLang 在高并发和长上下文场景下大幅提升了吞吐量,代价是 top-k 缓存未命中带来的 I/O 开销。

Sparse attention reduces KV cache memory and bandwidth requirements at the SDPA operation, but it does not reduce the overall memory capacity usage. In addition, HiSparse shows that system optimizations can overcome sparse attention’s memory capacity limitations, so sparse attention memory profile alone cannot sufficiently portray the full picture of system KV cache efficiency. To provide a more holistic picture of serving sparse attention models, here we explain the design of Z.ai’s GLM-5 model series, and how they are served in practice.

稀疏注意力减少了 SDPA 操作期间的 KV cache 内存和带宽需求,但并未减少整体内存容量的使用。此外,HiSparse 表明系统优化可以克服稀疏注意力的内存容量限制,因此仅凭稀疏注意力的内存配置不足以全面描绘系统 KV cache 效率的全貌。为了更完整地展示服务稀疏注意力模型的情况,此处我们介绍 Z.ai 的 GLM-5 模型系列的设计及其在实际中的服务方式。

GLM5.3 Agentic Inference Serving

GLM5.3 智能体推理服务

In our realtime benchmark board InferenceX, GB300 delivers the lowest modeled serving cost in our comparison at a response speed of 150 tokens per second. GB300 handles more tokens per GPU, and its higher hourly cost still leaves it ahead of GB200 on cost efficiency. The GB300 result uses Dynamo-TRT-LLM, while the GB200 result uses Dynamo-SGLang.

在我们实时基准测试平台 InferenceX 上,GB300 以每秒 150 个 token 的响应速度,在我们的对比中提供了最低的建模服务成本。GB300 每 GPU 处理的 token 更多,尽管其每小时成本更高,但在成本效益方面仍优于 GB200。GB300 的结果使用了 Dynamo-TRT-LLM,而 GB200 的结果使用了 Dynamo-SGLang。

We use the September 17 AgentX snapshot results to help illustrate the serving implications of GLM-5.3. We estimate costs using InferenceX’s model of owning and operating infrastructure at hyperscaler scale. Total tokens include input and generated output, including cached input history reused across turns. Streaming speed is measured using p90 interactivity; time to first token is evaluated separately.

我们使用9月17日AgentX快照结果来辅助说明GLM-5.3的服务影响。我们使用InferenceX的超大规模基础设施拥有和运营模型来估算成本。总token数包括输入和生成的输出,其中包括跨轮次重用的缓存输入历史。流式传输速度使用p90交互性进行测量;首token时间单独评估。

At 150 tokens per second, GB200 costs approximately $0.044 per million total tokens, compared with $0.238 for MI355X running ATOM. That is about 82% lower cost at the same streaming-speed target. These values are estimated from the measured curves, rather than from a separate test at exactly 150 tokens per second.

在每秒150个token的速度下,GB200每百万总token的成本约为0.044美元,而运行ATOM的MI355X为0.238美元。这比相同流式传输速度目标下的成本低约82%。这些值是从测量的曲线中估算得出的,而不是从恰好每秒150个token的单独测试中得出的。

For this comparison, we use InferenceX’s estimated cost of owning and operating the infrastructure at hyperscaler scale. Total tokens include input and generated output, including input history reused from cache.

对于此比较,我们使用InferenceX对超大规模基础设施拥有和运营成本的估算。总token数包括输入和生成的输出,其中包括从缓存中重用的输入历史。

Source: SemiAnalysis InferenceX current InferenceX dashboard; interpolation method; GB300 source run; MI355X source run.

来源:SemiAnalysis InferenceX当前InferenceX仪表板;插值方法;GB300源运行;MI355X源运行。

Among the Dynamo-SGLang results, GB300 serves roughly 13950 total tokens per second per GPU here, a 17.5% throughput advantage versus 11873 total tokens for GB200. However, we assume $2.31 per GB300 GPU hour versus $1.86 for GB200. The higher hourly cost more than offsets GB300’s throughput lead at this target.

在Dynamo-SGLang的结果中,GB300在此处每GPU每秒服务大约13950个总token,相比GB200的11873个总token具有17.5%的吞吐量优势。然而,我们假设GB300 GPU每小时成本为2.31美元,而GB200为1.86美元。更高的每小时成本完全抵消了GB300在此目标下的吞吐量领先优势。

In the comparison with AMD, GB200 costs about 52% less than ATOM at 100 tokens per second, 62% less at 125, and 82% less at 150. Across these three targets, the gap widens as response speed increases.

在与AMD的比较中,GB200在每秒100个token时比ATOM成本低约52%,在125时低62%,在150时低82%。在这三个目标中,随着响应速度的提高,差距扩大。

Counting only generated output, GB200 costs $5.92 per million output tokens at a response speed of 150 tokens per second, which is 79% less than $28.52 for MI355X ATOM. This metric is useful to analyze cost under agentic workloads, where agents repeatedly reuse long input histories while generating relatively little new text.

仅计算生成的输出,GB200在每秒150个token的响应速度下,每百万输出token的成本为5.92美元,比MI355X ATOM的28.52美元低79%。该指标有助于分析智能体工作负载下的成本,其中智能体反复重用长输入历史,同时生成相对较少的文本。

Source: SemiAnalysis InferenceX; current InferenceX dashboard

来源:SemiAnalysis InferenceX;当前InferenceX仪表板

The GB200 measurements on either side of the 150 token per second target have p90 TTFT (time to first token) of 14-19 seconds, compared with 1.3-1.7 seconds for MI355X ATOM. Some of the GB200 lowest cost runs have much longer waits for the first token.

在每秒150个token目标两侧的GB200测量中,p90 TTFT(首token时间)为14-19秒,而MI355X ATOM为1.3-1.7秒。部分GB200最低成本运行的首token等待时间要长得多。

We therefore make a second comparison using only configurations that were actually tested and achieved at least 150 tokens per second with p90 TTFT below two seconds. The chart below shows B200 with Dynamo-SGLang at $0.0666 per million total tokens versus $0.1265 for MI355X with SGLang. B200 still costs about 47% less with the first-token limit.

因此,我们仅使用实际测试过且 p90 TTFT 低于两秒、吞吐量达到每秒至少 150 个 token 的配置进行第二次比较。下图显示,B200 搭配 Dynamo-SGLang 的总 token 成本为每百万 0.0666 美元,而 MI355X 搭配 SGLang 的成本为每百万 0.1265 美元。在首 token 限制下,B200 的成本仍低约 47%。

Source: SemiAnalysis InferenceX B200 run 442071; GB300 run 441597; MI355X run 442075

来源:SemiAnalysis InferenceX B200 run 442071;GB300 run 441597;MI355X run 442075

Optimizations

优化措施

In GLM’s 5.2 / 5.3 model, its attention architecture reduces memory demands, and long running agents still need to retain earlier conversation history. GLM reduces the cost of handling this history in two ways: KV compression reduces the cached state stored per token, and sparse attention reduces how much of that state each attention operation reads. Then the serving engine would determine where to store the cache and how to retrieve it efficiently.

在 GLM 的 5.2 / 5.3 模型中,其注意力架构降低了对内存的需求,而长时间运行的智能体仍需保留早期的对话历史。GLM 通过两种方式降低了处理这些历史的成本:KV 压缩减少了每个 token 存储的缓存状态量,稀疏注意力减少了每次注意力操作读取的状态量。随后,推理引擎将决定如何存储缓存以及如何高效地检索它。

These B200 results show how much reuse can happen outside of GPU memory. When concurrency increases from 8 to 16 requests, the share of prompt tokens reused from GPU memory falls from 90.3% to 54.8%. Much of that reuse shifts to host memory, which rose from 6.0% to 40.3%. Much of the decline in GPU cache is offset by reuse from host memory, keeping the overall cache hit rate above 95% at all concurrency levels.

这些 B200 的结果展示了 GPU 内存之外可以发生多少复用。当并发请求数从 8 增加到 16 时,从 GPU 内存复用的 prompt token 比例从 90.3% 下降到 54.8%。大部分复用转移到了主机内存,后者从 6.0% 上升到 40.3%。GPU 缓存命中率的下降很大程度上被主机内存的复用所抵消,使得在所有并发级别下整体缓存命中率保持在 95% 以上。

Source: SemiAnalysis InferenceX; InferenceX B200 rows 440962, 440961, and 440959; workflow run 33683520699

来源:SemiAnalysis InferenceX;InferenceX B200 rows 440962, 440961, and 440959;workflow run 33683520699

Beyond cache management, changes to how the serving engine executes prefill and decoding can also change performance. Two separate GLM-5.2 studies illustrate this:

除了缓存管理之外,推理引擎执行预填充(prefill)和解码(decoding)方式的改变也会影响性能。两项独立的 GLM-5.2 研究说明了这一点:

  • vLLM keeping the first decoding step on a consistent CUDA graph execution path reduced average TPOT (time per output token) from about 40 ms to 22 ms in an NVFP4 study.
  • On a different 8 MI355X long context workload, ATOM engine processing prefill in chunks across pipeline stages delivered 98% higher total throughput and reduced median TTFT from 28.6 to 8.7 seconds.
  • 在一项 NVFP4 研究中,vLLM 将第一个解码步骤保持在一致的 CUDA 图执行路径上,使平均 TPOT(每个输出 token 的时间)从约 40 毫秒降低到 22 毫秒。
  • 在另一项涉及 8 个 MI355X 的长上下文工作负载中,ATOM 引擎跨流水线阶段以分块方式处理预填充,使总吞吐量提高了 98%,并将中位 TTFT 从 28.6 秒降低到 8.7 秒。

TileRT

For GLM5.3 on TileRT, AMD MI355X was supported first. TileRT is a low-latency LLM inference engine for that compiles decoding into a single persistent kernel, reducing launch overhead and overlapping computation, memory access, and communication.

对于 TileRT 上的 GLM5.3,AMD MI355X 率先得到支持。TileRT 是一个低延迟的大语言模型推理引擎,它将解码编译为单个持久化内核,从而减少启动开销,并实现计算、内存访问和通信的重叠。

It prioritizes faster token generation per user, rather than maximum batched throughput, while vLLM handle prefill.

它优先考虑每位用户更快的 token 生成速度,而非最大批量吞吐量,同时由 vLLM 处理预填充。

On AgentX, FP8 TileRT MI355X achieves 2x the P90 interactivity than the best FP4 MI335X config. Comparing to GB300 NVL72, it brings a 40% uplift in interactivity. These are serving metrics that used to require specialized hardware.

在 AgentX 上,FP8 TileRT MI355X 的 P90 交互性比最佳 FP4 MI335X 配置高出 2 倍。与 GB300 NVL72 相比,其交互性提升了 40%。这些是以往需要专用硬件才能实现的推理指标。

However, TTFT is suboptimal and there is still room for development, such as optimizing KV transfer, FP4 support, and larger batch sizes.

然而,TTFT(首 token 延迟)并非最优,仍有改进空间,例如优化 KV 传输、支持 FP4 以及增大批处理大小。

DeepSeek Sparse Attention

DeepSeek 稀疏注意力机制

GLM-5 (and GLM-5.x models) is a 744B total, 40B active-parameter mixture-of-experts model. For every token, it has 1 shared expert and is routed through 8 out of 256 experts, which is sparsity 32. It features DeepSeek Sparse Attention, which we will discuss in this section.

GLM-5(及 GLM-5.x 系列模型)是一个总参数量 744B、激活参数 40B 的混合专家(MoE)模型。对于每个 token,它包含一个共享专家,并从 256 个专家中路由至 8 个专家,稀疏度为 32。该模型具备 DeepSeek 稀疏注意力机制,我们将在本节中进行讨论。

DeepSeek Sparse Attention (DSA), introduced in DeepSeek V3.2, consists of two components: A lightning indexer that selects top K tokens, and a sparse Multi-Latent Attention (MLA).

DeepSeek 稀疏注意力机制(DSA)于 DeepSeek V3.2 中引入,由两个组件构成:一个用于选择 Top K token 的快速索引器(Lightning Indexer),以及一个稀疏的多潜注意力(Sparse Multi-Latent Attention, MLA)。

Lightning Indexer

快速索引器

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件