跳到主内容
@wquguru
精选88NVIDIA 博客(RSS)行业动态

NVIDIA Vera Rubin NVL72在MLPerf

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

原文
发到 X
推荐理由

Vera Rubin作为下一代架构的首批实测数据,直接反映算力代际跃迁,对评估推理成本与选型极具参考价值。

System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments.

系统性能、高效的基础设施扩展以及持续的软件优化是决定 AI 推理经济性的关键杠杆。更高的系统性能意味着生成更多的 token,从而带来更高的收入。高效的扩展意味着随着硬件的增加,吞吐量成比例增长,从而以较少的资源服务大规模用户。持续的优化意味着从基础设施投资中产生更多的价值。

Underlying all three is platform fungibility: the same infrastructure runs any model, any workload, from training to inference, recommender to reasoning, language to video, keeping utilization high.

在这三者背后的基础是平台的可互换性:相同的基础设施可以运行任何模型、任何工作负载,涵盖从训练到推理、推荐器到推理引擎、语言到视频的所有领域,从而保持高利用率。

The NVIDIA platform is purpose-built to optimize across all these, as highlighted by MLPerf Inference v6.1 results released today:

NVIDIA 平台专为在所有这些方面进行优化而构建,正如今天发布的 MLPerf Inference v6.1 结果所强调的那样:

  • NVIDIA Vera Rubin NVL72 system debuts with leading performance: In its first MLPerf Inference preview submission, NVIDIA Vera Rubin NVL72 delivers up to 3.7x better throughput than GB300 NVL72.
  • NVIDIA GB300 NVL72 scales with leading efficiency: A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, with throughput growing nearly linearly from a single-rack baseline.
  • Continuous software optimizations drive performance gains: Software optimizations in NVIDIA’s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance over v6.0. Optimizations continued post-v6.1 submission, delivering further performance gains.
  • NVIDIA Vera Rubin NVL72 系统以领先的性能首次亮相:在首次 MLPerf Inference 预览提交中,NVIDIA Vera Rubin NVL72 的吞吐量比 GB300 NVL72 高出高达 3.7 倍。
  • NVIDIA GB300 NVL72 以领先的效率实现扩展:跨四个 GB300 NVL72 机架的 288 GPU 提交实现了 99% 的扩展效率,吞吐量从单机架基线几乎呈线性增长。
  • 持续的软件优化推动性能提升:NVIDIA 在 MLPerf Inference v6.1 提交中的软件优化带来了比 v6.0 高出高达 1.6 倍的性能。在 v6.1 提交之后,优化仍在继续,带来了进一步的性能提升。

For organizations making AI infrastructure decisions, performance, scaling efficiency and software velocity are important considerations that determine long-term inference economics.

对于做出 AI 基础设施决策的组织而言,性能、扩展效率和软件速度是重要的考量因素,它们决定了长期的推理经济性。

Vera Rubin NVL72 Makes MLPerf Inference Debut With Leading Performance

Vera Rubin NVL72 以领先的性能首次亮相 MLPerf Inference

NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL.

NVIDIA 提交了 Vera Rubin NVL72 的预览结果,针对 MLPerf Inference v6.1 套件中最具挑战性的两个基准测试:DeepSeek-R1 和 Qwen3-VL。

Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios, using vLLM with the NVIDIA Dynamo open source inference framework. On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, throughput is up to 2.5x higher than GB300 NVL72. These early results showcase NVIDIA’s accelerated pace of innovation and how performance will improve with continuous software optimizations.

在离线、服务器和交互场景下,使用带有 NVIDIA Dynamo 开源推理框架的 vLLM,Vera Rubin NVL72 在 Qwen3-VL 上的吞吐量比 GB300 NVL72 高出高达 3.7 倍。在使用 NVIDIA TensorRT-LLM 库的 DeepSeek-R1 上,吞吐量比 GB300 NVL72 高出高达 2.5 倍。这些早期结果展示了 NVIDIA 加速的创新步伐,以及持续软件优化将如何进一步提升性能。

MLPerf Inference v6.1, Closed Division. Results retrieved from www.mlcommons.org on Sep 16, 2026. NVIDIA platform results from the following entries: 6.1-0106 and 6.1-0074. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information.

MLPerf 推理 v6.1,封闭组。结果于 2026 年 9 月 16 日从 www.mlcommons.org 获取。NVIDIA 平台结果来自以下条目:6.1-0106 和 6.1-0074。MLPerf 名称和标志是 MLCommons Association 在美国及其他国家的注册和非注册商标。保留所有权利。严禁未经授权使用。有关更多信息,请访问 www.mlcommons.org。

This performance means each Vera Rubin NVL72 rack delivers significantly more tokens, serves more users and generates more revenue than a GB300 NVL72 rack, while lowering cost per token.

这一性能意味着每个 Vera Rubin NVL72 机架相比 GB300 NVL72 机架能交付显著更多的 token、服务更多用户并产生更高收入,同时降低每个 token 的成本。

The results reflect full-stack codesign across hardware and software. Vera Rubin’s enhanced Tensor Cores and Transformer Engine accelerate both the prefill and decode stages of inference, while NVFP4 precision reduces memory footprint across model weights, attention and KV cache — increasing throughput with minimal loss of output quality.

这些结果反映了硬件与软件的全栈协同设计。Vera Rubin 的增强型 Tensor Core 和 Transformer Engine 加速了推理的预填充(prefill)和解码(decode)阶段,而 NVFP4 精度降低了模型权重、注意力机制和 KV 缓存的内存占用——在输出质量损失极小的情况下提高了吞吐量。

Vera Rubin submissions heavily used disaggregated serving, separating prefill and decode along with large-scale expert parallelism for maximum efficiency across the mixture-of-experts layers that power models like DeepSeek-R1 and Qwen3-VL.

Vera Rubin 的提交大量使用了分离式推理服务(disaggregated serving),将预填充和解码分离,并结合大规模专家并行技术,以最大化像 DeepSeek-R1 和 Qwen3-VL 这类模型所依赖的混合专家(MoE)层的效率。

The NVL72 scale-up domain — powered by sixth-generation NVIDIA NVLink and NVLink Switch to deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet — provides the interconnect foundation that makes these techniques effective at rack scale.

NVL72 规模扩展域——由第六代 NVIDIA NVLink 和 NVLink Switch 驱动,提供比现成以太网高 10 倍的包速率和低 3 倍的延迟——提供了使这些技术在机架规模上有效的互连基础。

This codesign extends to NVIDIA’s partner ecosystem: Nebius also submitted Vera Rubin NVL72 preview results and demonstrated excellent performance.

这种协同设计延伸至 NVIDIA 的合作伙伴生态系统:Nebius 也提交了 Vera Rubin NVL72 的预览结果,并展示了卓越的性能。

AI agents, which reason, plan and act across multiple steps, are reshaping how inference performance is measured. In benchmarks designed to capture this shift, such as SemiAnalysis AgentX, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing. In addition, the upcoming MLPerf Endpoints benchmark will bring standardized measurement to agentic inference workloads, beyond what traditional throughput benchmarks capture.

AI 智能体能够在多个步骤中进行推理、规划和行动,正在重塑推理性能的衡量方式。在旨在捕捉这一转变的基准测试中,例如 SemiAnalysis AgentX,Vera Rubin NVL72 在预览测试中比 GB300 NVL72 提供了 30 倍更好的性能。此外,即将推出的 MLPerf Endpoints 基准测试将为智能体推理工作负载带来标准化测量,超越传统吞吐量基准所能捕捉的范围。

NVIDIA GB300 NVL72 Scales With Leading Efficiency

NVIDIA GB300 NVL72 以领先的效率实现扩展

Scaling efficiency — how effectively additional GPUs translate to throughput gains — is a key measure of AI infrastructure productivity. NVIDIA delivers this with high-bandwidth, low-latency scale-up interconnects within each rack, high-bandwidth networking between racks and efficient request orchestration across nodes.

扩展效率——即额外 GPU 转化为吞吐量提升的有效性——是衡量 AI 基础设施生产力的关键指标。NVIDIA 通过机架内的高带宽、低延迟规模扩展互连、机架间的高带宽网络以及跨节点的高效请求编排来实现这一点。

NVIDIA’s DeepSeek-R1 (DSR1) submission scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs), achieving 99% scaling efficiency in the offline scenario. Throughput grew nearly in proportion to the hardware added.

NVIDIA 的 DeepSeek-R1 (DSR1) 提交从单个 GB300 NVL72 机架(72 个 GPU)扩展到四个机架(288 个 GPU),在离线场景中实现了 99% 的扩展效率。吞吐量随硬件增加几乎呈线性增长。

MLPerf Inference v6.1, Closed Division. Results retrieved from www.mlcommons.org on Sep 16, 2026. NVIDIA platform results from the following entries: 6.1-0073 and 6.1-0074. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information.

MLPerf Inference v6.1,封闭组。结果于 2026 年 9 月 16 日从 www.mlcommons.org 获取。NVIDIA 平台结果来自以下条目:6.1-0073 和 6.1-0074。MLPerf 名称和标志是 MLCommons Association 在美国及其他国家的注册和非注册商标。保留所有权利。严禁未经授权使用。有关更多信息,请访问 www.mlcommons.org。

Scaling efficiency is key because more GPUs don’t automatically mean proportionally more throughput. If adding nearly double the GPU count delivered only a single-digit percentage improvement in throughput, the infrastructure cost would far outpace the performance return. The architecture, interconnect and software must all scale together.

扩展效率至关重要,因为更多的 GPU 并不自动意味着成比例的更高吞吐量。如果增加近一倍的 GPU 数量仅带来个位数的吞吐量提升,基础设施成本将远远超过性能回报。架构、互连和软件必须协同扩展。

GB300 NVL72 also demonstrated rack-scale efficiency on the WAN 2.2 text-to-video benchmark, reaching 0.65 720p videos per second at 5.7 seconds per video — 9x higher throughput and 7.5x lower latency than a single node.

GB300 NVL72 在 WAN 2.2 文本到视频基准测试中也展示了机架级效率,达到每秒 0.65 个 720p 视频,每个视频耗时 5.7 秒——吞吐量比单节点高 9 倍,延迟低 7.5 倍。

Software Optimizations Drive Continuous Gains

软件优化驱动持续增益

NVIDIA platform undergoes continuous software development, delivering performance and feature improvements.

NVIDIA 平台持续进行软件开发,提供性能和功能改进。

In v6.1, GB300 NVL72 performance on Qwen3-VL improved up to 1.6x over v6.0 results. The gains came through lower KV cache precision, additional kernel fusion, better kernels and disaggregated serving with vLLM and NVIDIA Dynamo.

在 v6.1 中,GB300 NVL72 在 Qwen3-VL 上的性能相比 v6.0 结果最高提升了 1.6 倍。这些增益通过降低 KV cache 精度、额外的内核融合、更好的内核以及使用 vLLM 和 NVIDIA Dynamo 进行的解耦服务实现。

Software optimization continued past the v6.1 submission deadline as well. Post-submission results, not yet verified by MLCommons, on GPT-OSS-120B and DLRMv3 show further performance gains.

软件优化在 v6.1 提交截止日期之后仍在继续。提交后结果(尚未由 MLCommons 验证)显示,GPT-OSS-120B 和 DLRMv3 的性能进一步提升。

AI Inference at Every Scale

全规模 AI 推理

Beyond the NVIDIA Grace Blackwell and Vera Rubin NVL72 platform results, NVIDIA submitted Jetson AGX Thor results using NVIDIA TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B.

除了 NVIDIA Grace Blackwell 和 Vera Rubin NVL72 平台的结果外,NVIDIA 还提交了使用 NVIDIA TensorRT Edge-LLM 在新推出的 Edge-Agentic 基准测试(配合 Qwen3.6-27B)中的 Jetson AGX Thor 结果。

The NVIDIA partner ecosystem participated broadly, with 19 partners — eight of them on multi-node Blackwell NVL72 systems — demonstrating excellent performance. This includes ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, HPE, Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro and Wiwynn.

NVIDIA 合作伙伴生态系统广泛参与,19 家合作伙伴(其中八家基于多节点 Blackwell NVL72 系统)展示了卓越的性能。其中包括 ASUS、Azure、Cisco、CoreWeave、Crusoe、Dell Technologies、Fujitsu、Giga Computing、HPE、Inventec、Lambda、MiTAC Computing、Nebius、Oracle Cloud Infrastructure、Quanta Cloud Technology、Red Hat、ScitiX、Supermicro 和 Wiwynn。

From compact edge devices to the largest AI factories, NVIDIA continues to advance performance across the full technology stack with an annual cadence of platform architectures, continuously improving software and an ecosystem built to deliver it at scale.

从紧凑型边缘设备到最大的 AI 工厂,NVIDIA 持续推动整个技术栈的性能提升,以年度节奏推出平台架构,不断优化软件以及旨在大规模交付的生态系统。

Learn more about the NVIDIA Vera Rubin platform.

了解更多关于 NVIDIA Vera Rubin 平台的信息。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件