跳到主内容
@wquguru
精选86NVIDIA 博客(RSS)产品发布/更新

NVIDIA DSX平台实测:固定功耗下算力提升24%

From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

原文
发到 X
推荐理由

这是目前最硬核的AI基建效率优化方案,Lambda实测数据直接证明了DSX能在不扩容电网的情况下显著提升算力产出,做Infra的同学值得重点关注其落地效果。

On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to adjust its power consumption.

在硅谷一个酷热难耐的八月傍晚,随着太阳西沉和空调负荷激增,硅谷电力公司向一家AI工厂发送信号,要求其调整用电量。

Varun Sivaram was watching on Zoom with about forty others — his team at Emerald AI in their San Francisco conference room, engineers at the data center and people from the utility itself. Nobody touched anything.

Varun Sivaram 正通过 Zoom 与大约四十人一起观看——他的 Emerald AI 团队位于旧金山的会议室里,数据中心工程师以及公用事业公司的人员也在其中。没有人动手操作任何东西。

Emerald AI’s Conductor platform — a grid-orchestration platform from NVIDIA partner Emerald AI, and an early example of the kind of flexibility NVIDIA DSX Flex is built to deliver — receives signals about grid conditions and adjusts the data center’s flexible computing workloads. Work that can wait is slowed or rescheduled, while higher-priority services continue operating.

Emerald AI 的 Conductor 平台——来自 NVIDIA 合作伙伴 Emerald AI 的电网编排平台,也是 NVIDIA DSX Flex 旨在提供的灵活性的早期示例——接收有关电网状况的信号,并调整数据中心的弹性计算工作负载。可以等待的工作会被减速或重新安排,而高优先级服务则继续运行。

The goal is to reduce electricity demand when the grid is constrained without interrupting critical AI workloads— exactly what Silicon Valley Power needed,

目标是在电网受限的情况下减少电力需求,同时不中断关键的 AI 工作负载——这正是硅谷电力公司所需要的,

When the reduction showed on screen, everyone cheered.

当屏幕上的减量显示出来时,所有人都欢呼起来。

“We were watching with bated breath,” Sivaram said. “It was our first time deploying across thousands of NVIDIA GPUs.” His head of product, Mansi Shah, was emotional. “This feels kind of like a SpaceX rocket launch,” she said.

“我们屏息观看,”Sivaram 说。“这是我们首次部署跨越数千个 NVIDIA GPU。”他的产品负责人 Mansi Shah 情绪激动。“这感觉有点像 SpaceX 的火箭发射,”她说。

Silicon Valley Power has since sent more than 200 demand signals to that AI factory. It worked every single time.

此后,硅谷电力公司已向该 AI 工厂发送了超过 200 次需求响应信号。每次都很成功。

The Emerald AI team in San Francisco watches as Silicon Valley Power’s demand signal hits the factory floor — power dropping from four megawatts to three, automatically, while every high-priority job keeps running.

旧金山的 Emerald AI 团队注视着硅谷电力的需求信号到达工厂车间——功率从四兆瓦自动降至三兆瓦,同时所有高优先级作业仍在运行。

This is grid flexibility in production. And it points at something much bigger than one facility in Santa Clara: a path to unlocking the power America’s AI factories need, without waiting a decade to build new transmission lines.

这是生产环境中的电网灵活性。它指向的远不止圣克拉拉的一个设施:这是一条释放美国 AI 工厂所需潜力的路径,无需等待十年去建设新的输电线路。

At the AI Infra Summit on Tuesday, Ian Buck, NVIDIA’s vice president of hyperscale and high-performance computing, made AI factory efficiency the centerpiece of his infrastructure keynote.

在周二举行的 AI 基础设施峰会上,NVIDIA 超大规模和高性能计算副总裁 Ian Buck 将 AI 工厂效率作为其基础设施主题演讲的核心。

Results from cloud provider Lambda’s first validation in a deployment environment, released the same day, put numbers to it: a fixed power budget can support 24% more token throughput when managed intelligently.

云服务商 Lambda 在同一天发布的其在部署环境中首次验证的结果给出了具体数据:在智能管理下,固定的电力预算可支持 24% 更高的 token 吞吐量。

“With our proof of concept, we believe we’ve moved beyond the limitation of fixed power budgets,” said Dave Ward, president of cloud services at Lambda. “NVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint.”

“通过我们的概念验证,我们相信已经突破了固定电力预算的限制,”Lambda 云服务总裁 Dave Ward 表示。“NVIDIA DSX MaxLPS 为回收闲置容量并将其转化为实际使用铺平了道路,在相同的占地面积内实现显著更高的计算密度。”

That August evening, when SVP called, Conductor executed against a predefined workload hierarchy: lowest-priority jobs yielded, high-priority inference kept running, and power fell from four megawatts to three. Automated. No operator required.

那年八月的一天晚上,当SVP来电时,Conductor按照预定义的工作负载层级执行:低优先级作业让路,高优先级推理继续运行,功率从四兆瓦降至三兆瓦。全自动,无需人工干预。

DSX at a glance

DSX概览

NVIDIA DSX MaxLPS

NVIDIA DSX MaxLPS

A suite of technologies to optimize AI factory throughput per megawatt, including dynamic power allocation software that monitors GPU and rack-level consumption in real time, recovering stranded capacity to maximize token throughput within a fixed power budget.

一套旨在优化每兆瓦AI工厂吞吐量的技术组合,包括动态功率分配软件,可实时监控GPU和机柜级别的功耗,回收闲置容量,从而在固定预算内最大化Token吞吐量。

NVIDIA DSX Flex

NVIDIA DSX Flex

Receives grid signals (load-shedding, demand-response, pricing events) and adapts AI workload priorities in response, protecting high-priority jobs while reducing overall power draw.

接收电网信号(如负荷削减、需求响应、价格事件)并据此调整AI工作负载优先级,在保护高优先级作业的同时降低整体功耗。

NVIDIA DSX OS

NVIDIA DSX OS

Open-source, modular software for AI factory lifecycle management, runtime consistency, health automation and resiliency.

开源、模块化的AI工厂生命周期管理软件,提供运行时一致性、健康自动化及弹性保障。

NVIDIA DSX Sim

NVIDIA DSX Sim

Simulation tools that let operators model and validate factory designs before physical deployment, identifying bottlenecks before capital is fixed.

仿真工具,使操作员能够在物理部署前对工厂设计进行建模和验证,在资本固化之前识别瓶颈。

NVIDIA DSX Reference Designs

NVIDIA DSX Reference Designs

Generation-specific, validated architectures spanning compute, networking, storage and facilities, co-designed with NVIDIA’s ecosystem partners.

针对特定代际的已验证架构,涵盖计算、网络、存储和设施,由NVIDIA与其生态合作伙伴共同设计。

In the AI factory economy, power is the constraint. Work per gigawatt is the metric. Data center operators are meticulous about efficiency — every watt put to work is a watt delivering productive compute, and the industry has driven remarkable gains at every layer of the stack, from facility design to rack-level power conversion.

在AI工厂经济中,电力是约束条件,每吉瓦的处理量是关键指标。数据中心运营商对效率精益求精——每一瓦特投入工作的电力都必须转化为有效的计算产出。整个行业在堆栈的每一层都取得了显著进步,从设施设计到机柜级电源转换。

DSX extends that discipline into the AI workload itself. Smarter rack provisioning puts power where workloads actually need it. Operational intelligence — tighter scheduling, faster restarts, leaner checkpointing — keeps GPUs running rather than waiting. The goal is the same one operators have always pursued: more work from the power you have.

DSX将这种严谨性延伸至AI工作负载本身。更智能的机柜配置将电力精准投放到工作负载真正需要的地方。运营智能——更紧密的调度、更快的重启、更精简的检查点机制——确保GPU持续运行而非等待。其目标与运营商一直追求的一致:用现有的电力完成更多工作。

“A one-gigawatt factory will never become a two-gigawatt factory,” NVIDIA founder and CEO Jensen Huang has said.

“一座一吉瓦的工厂永远不会变成两座吉瓦的工厂。” NVIDIA创始人兼CEO黄仁勋曾如是说。

The answer engineers reach when systems hit physical limits is always the same: stop optimizing the parts and start designing the whole.

当系统触及物理极限时,工程师得出的答案始终如一:停止优化局部部件,开始设计整体系统。

Introduced at GTC Taipei in May, NVIDIA DSX is that answer for the AI factory — and the early deployments are already proving it out.

今年五月在台北GTC大会上推出的NVIDIA DSX正是这一答案,且早期部署已证明其有效性。

The full platform spans networking, cooling, water efficiency and facility design; the sections below focus on some of the results so far in power management and grid participation.

该平台涵盖网络、冷却、水效和设施设计;下文部分将聚焦于目前在电力管理和电网参与方面取得的部分成果。

More Compute, Same Budget: DSX MaxLPS

更多算力,相同预算:DSX MaxLPS

Lambda’s results, released at the AI Infra Summit, are the first validation of DSX MaxLPS on NVIDIA HGX B200 GPU Servers.

Lambda 在 AI 基础设施峰会上发布的成果,是 DSX MaxLPS 在 NVIDIA HGX B200 GPU 服务器上的首次验证。

DSX MaxLPS monitors GPU and rack-level power consumption and reallocates headroom across nodes based on workload type, recovering capacity that static provisioning would leave stranded. Training and inference draw power differently; MaxLPS optimizes allocation in AI factories running both.

DSX MaxLPS 监控 GPU 和机柜级别的功耗,并根据工作负载类型在各节点间重新分配余量,从而回收静态配置会遗留的闲置容量。训练和推理的功耗特性不同;MaxLPS 优化了同时运行这两类任务的 AI 工厂中的资源分配。

Lambda, a GPU cloud provider serving more than 10,000 customers from AI-native startups to hyperscalers, ran the software on a five-rack, 19-node cluster.

Lambda 是一家 GPU 云服务商,客户超过 10,000 家,涵盖从原生 AI 初创公司到超大规模企业。该软件在其由五个机柜、19 个节点组成的集群上进行了运行测试。

What they found: by running 19 nodes within the same power budget as 16 nodes at full power, Lambda achieved 24% more cluster-wide token throughput — from roughly 4 million tokens per second to 5 million. Performance per watt improved by 23%.

他们的发现:在仅相当于 16 个节点满功率运行的同一电力预算内运行 19 个节点,Lambda 实现了集群整体 token 吞吐量提升 24%——从每秒约 400 万 token 提升至 500 万。每瓦性能提升了 23%。

Based on NVIDIA’s projections, DSX MaxLPS can enable up to 40% more GPU capacity for next-generation Vera Rubin NVL72 AI factories within the same megawatt power budget in suitable deployment environments.

根据 NVIDIA 的预测,在合适的部署环境中,DSX MaxLPS 可在相同的兆瓦级电力预算内,为下一代 Vera Rubin NVL72 AI 工厂提供多达 40% 的额外 GPU 容量。

Automated Demand Response, Proven in Production

自动化需求响应,已在生产中验证

The Santa Clara story isn’t a DSX Flex installation — it’s something earlier and more important: proof that the concept works at commercial scale.<

圣克拉拉的故事并非 DSX Flex 的安装案例——而是更早且更重要的内容:证明了该概念在商业规模下可行。<

NVIDIA’s Eos AI factory is running Emerald AI Conductor as a participant in Silicon Valley Power‘s Flexible Load Interconnect Program, the first commercial grid utility program designed to treat AI factories as dispatchable resources.

NVIDIA 的 Eos AI 工厂正在作为参与者运行 Emerald AI Conductor,加入硅谷电力(Silicon Valley Power)的可调负荷互联计划(Flexible Load Interconnect Program),这是首个旨在将 AI 工厂视为可调度资源的商业电网公用事业项目。

When Silicon Valley Power sends a signal, Conductor responds in under a minute. The factory that’s willing to flex gets to run bigger.

当硅谷电力发送信号时,Conductor 能在不到一分钟的时间内做出响应。愿意进行负荷调节的工厂得以扩大运行规模。

That’s the pattern DSX Flex is built to generalize — with Emerald AI Conductor integrating into DSX Flex as the platform matures. The first dedicated DSX Flex commercial deployment will be the Manassas, Virginia, facility: a 96-megawatt Vera Rubin AI factory at NVIDIA’s AI Factory Research Center, building on five prior demonstrations across two continents.

这正是 DSX Flex 旨在泛化的模式——随着平台成熟,Emerald AI Conductor 将集成到 DSX Flex 中。首个专用的 DSX Flex 商业部署将是位于弗吉尼亚州马纳萨斯的设施:这是 NVIDIA AI 工厂研究中心的一个 96 兆瓦 Vera Rubin AI 工厂,建立在跨越两大洲的五次先前演示基础之上。

The Next Power Architecture Layer: 800V DC Power Architecture

下一个电源架构层:800V 直流电源架构

The gains inside today’s AI factory are real and deployable now. The next layer is how power is delivered to denser accelerated computing racks.

当前 AI 工厂内部的增益是真实存在的,且现已可部署。下一层涉及如何向更密集的加速计算机柜供电。

As AI factories scale, traditional lower-voltage power paths add conversion complexity and distribution constraints.

随着 AI 工厂规模的扩大,传统的低压供电路径增加了转换复杂性和配电限制。

NVIDIA’s 800 VDC architecture is designed to reduce conversion complexity, improve power delivery efficiency and support denser accelerated computing racks.

NVIDIA 的 800 VDC 架构旨在降低转换复杂性,提高供电效率,并支持更密集的加速计算机柜。

NVIDIA DSX is incorporating 800V DC into its reference designs.

NVIDIA DSX 正在将其参考设计纳入 800V 直流技术。

The Whole Factory, Not the Parts

整个工厂,而非零部件

No single component can optimize an AI factory on its own. A faster GPU still waits on the network. Power can be stranded by bad provisioning. Cooling overhead still diverts electricity from GPUs; GB200 NVL72 racks running direct liquid cooling carry ~120 kW of heat that has to go somewhere before that power reaches compute.

没有任何单一组件能够独立优化 AI 工厂。更快的 GPU 仍会受制于网络瓶颈。供电可能因配置不当而闲置。散热开销仍在分流本可用于计算的电力;采用直接液冷的 GB200 NVL72 机架会产生约 120 kW 的热量,这些热量必须在电力转化为计算能力之前被导出。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件