跳到主内容
@wquguru
精选75NVIDIA 博客(RSS)行业动态

NVLink Fusion 发布:让定制 XPU 接入英伟达 AI 工厂基础设施

How XPUs Meet a World-Class AI Factory

原文
发到 X

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime.

为了大规模生成智能,AI工厂持续运行,其经济性由交付产出定义:每秒令牌数、每瓦特令牌数、每令牌成本、利用率和正常运行时间。

That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators.

这要求AI基础设施作为完整的工厂来设计和构建,而非单个加速器的集合。

Hyperscalers and AI-native companies building custom XPUs must consider not just XPU design, but the design and development of the entire AI platform, including scale-up and scale-out networking, rack-scale architecture, production factory software and a robust supplier ecosystem.

超大规模企业和AI原生公司构建定制XPU时,不仅需要考虑XPU设计,还需考虑整个AI平台的设计与开发,包括纵向扩展和横向扩展网络、机架级架构、生产工厂软件以及强大的供应商生态系统。

At AI factory scale, this path is complex and costly, and represents a fundamental obstacle to getting XPUs to market quickly.

在AI工厂规模下,这条路径复杂且成本高昂,是XPU快速上市的根本障碍。

Breaking the constraint means combining custom XPUs with proven, mature infrastructure — allowing builders to focus innovation where it matters most while harnessing established technology for the rest.

打破这一限制意味着将定制XPU与成熟、经过验证的基础设施相结合——让构建者将创新集中在最关键的地方,同时利用成熟技术处理其余部分。

NVLink Fusion delivers on that need, connecting XPUs to NVIDIA’s world-leading AI infrastructure to increase performance, accelerate time to market and mitigate risk for semi-custom AI factories.

NVLink Fusion满足了这一需求,将XPU连接到NVIDIA世界领先的AI基础设施,以提高性能、加速上市时间,并降低半定制AI工厂的风险。

Unlock XPU Performance With Fast Scale-Up

通过快速纵向扩展释放XPU性能

For modern workloads such as running trillion-parameter models, mixture-of-experts architectures and agentic AI, if the scale-up fabric cannot keep up, utilization drops and cost per token rises.

对于现代工作负载,如运行万亿参数模型、混合专家架构和代理式AI,如果纵向扩展网络无法跟上,利用率就会下降,每令牌成本上升。

A scale-up networking solution must excel on three dimensions:

纵向扩展网络解决方案必须在三个维度上表现出色:

  • Delivered performance: End-to-end network performance, in-network compute and mature software integration.
  • Factory resiliency: Uptime, continuous health monitoring and telemetry, and component-level serviceability while the factory keeps running.
  • Platform maturity: Reduced operational risk by using a mature technology stack with a demonstrated track record of large-scale deployments and realized return on investment.
  • 交付性能:端到端网络性能、网络内计算和成熟的软件集成。
  • 工厂弹性:正常运行时间、持续健康监控和遥测,以及在工厂持续运行时的组件级可维护性。
  • 平台成熟度:通过使用具有大规模部署和实现投资回报记录的成熟技术栈,降低运营风险。

As an example, NVLink Fusion brings XPUs into the NVIDIA NVLink scale-up domain. Sixth-generation NVLink provides leading high-bandwidth, low-latency networking across a 72-XPU domain. The end-to-end latency for XPU-to-XPU transfers is 3x lower than alternative solutions based on off-the-shelf Ethernet, and the packet rate is 10x higher.

例如,NVLink Fusion将XPU引入NVIDIA NVLink纵向扩展域。第六代NVLink在72-XPU域内提供领先的高带宽、低延迟网络。XPU到XPU传输的端到端延迟比基于现成以太网的替代方案低3倍,数据包速率高10倍。

For end-to-end performance, NVIDIA GB300 NVL72 systems help deliver significantly higher throughput and better interactivity compared with configurations that don’t use NVL72, and future NVLink roadmap configurations include domains of up to 1,152 accelerators and co-packaged optics.

对于端到端性能,与不使用NVL72的配置相比,NVIDIA GB300 NVL72系统有助于显著提高吞吐量和更好的交互性,未来的NVLink路线图配置包括多达1,152个加速器的域和共封装光学器件。

The 72-GPU NVLink scale-up domain enables GB300 NVL72 to deliver higher per-GPU throughput and interactivity compared with NVIDIA B300. Results from NVIDIA’s AI Inference Performance Benchmarks page.

72-GPU NVLink 扩展域使 GB300 NVL72 相比 NVIDIA B300 能够提供更高的每 GPU 吞吐量和交互性。结果来自 NVIDIA AI 推理性能基准页面。

NVLink Fusion also includes NVIDIA NVLink-C2C for connecting XPUs to NVIDIA Vera CPUs or other ecosystem CPUs, delivering up to 6x the energy efficiency of a PCIe interface — helping remove barriers between control and compute for agentic systems.

NVLink Fusion 还包括 NVIDIA NVLink-C2C,用于将 XPU 连接到 NVIDIA Vera CPU 或其他生态系统 CPU,能效比 PCIe 接口高出 6 倍——有助于消除智能体系统中控制与计算之间的障碍。

A Proven Stack and Ecosystem for Development and Deployment

成熟的软件栈和生态系统,支持开发与部署

Teams developing custom XPUs often underestimate the effort and complexity of turning XPU innovation into data center deployment. This includes:

开发定制 XPU 的团队往往低估了将 XPU 创新转化为数据中心部署所需的努力和复杂性。这包括:

  • Integrating high-speed CPU and scale-up interfaces
  • Sourcing and validating a scale-up network solution
  • Designing compute and switch trays
  • Designing and validating a rack architecture, including cooling and power
  • Integrating security and storage
  • Managing a complex supplier ecosystem
  • 集成高速 CPU 和扩展接口
  • 采购并验证扩展网络解决方案
  • 设计计算和交换托盘
  • 设计并验证机架架构,包括冷却和供电
  • 集成安全和存储
  • 管理复杂的供应商生态系统

The ideal platform provides all of this, allowing teams to focus on targeted innovation while using proven solutions for the rest.

理想的平台应提供所有这些,使团队能够专注于有针对性的创新,同时使用经过验证的解决方案来处理其余部分。

NVLink Fusion is supported by an ecosystem designed for rapid development, integration and deployment, spanning ASIC design, CPU, and IP and optical interconnect partners.

NVLink Fusion 得到了一个专为快速开发、集成和部署而设计的生态系统的支持,涵盖 ASIC 设计、CPU、IP 和光互连合作伙伴。

“NVLink Fusion gives customers the ability to choose the CPU architecture, the performance level, the software capabilities that best meet their needs for the workloads that they care about,” said Tim Wilson, vice president and general manager of data center silicon engineering at Intel.

“NVLink Fusion 让客户能够选择最适合其工作负载需求的 CPU 架构、性能水平和软件功能,”英特尔数据中心硅工程副总裁兼总经理 Tim Wilson 表示。

NVLink Fusion adopters can also use the NVIDIA MGX rack-scale architecture and the same supply chain used for MGX-based systems such as NVIDIA Vera Rubin NVL72. Manufacturing partners manage design and integration, while MGX suppliers provide the building blocks for rack, cooling, power and emerging 800 VDC designs.

NVLink Fusion 的采用者还可以使用 NVIDIA MGX 机架级架构,以及用于基于 MGX 的系统(如 NVIDIA Vera Rubin NVL72)的相同供应链。制造合作伙伴负责设计和集成,而 MGX 供应商提供机架、冷却、供电和新兴 800 VDC 设计的构建模块。

“With Vera Rubin [NVL72], we are looking at almost 100% automation of system builds in the manufacturing line,” said Jack Luoh, head of product and solution at QCT and Quanta Computer. “Most of those investments can be leveraged if the XPU leverages NVLink Fusion.”

“借助 Vera Rubin [NVL72],我们正在实现生产线中系统构建的几乎 100% 自动化,”QCT 和广达电脑产品与解决方案负责人 Jack Luoh 表示。“如果 XPU 采用 NVLink Fusion,这些投资中的大部分都可以得到利用。”

Managing Risk With Infrastructure Standardization

通过基础设施标准化管理风险

AI factory planning doesn’t wait for silicon. Power procurement, facility design, cooling, rack layout and network architecture begin long before the final accelerator mix is available. A data center locked to one chip can become a schedule risk.

AI 工厂的规划不会等待芯片。电力采购、设施设计、冷却、机架布局和网络架构在最终加速器组合确定之前很久就开始了。锁定单一芯片的数据中心可能成为进度风险。

Different workloads may favor different accelerators, including XPUs, GPUs, CPUs and LPUs. GPU systems may work alongside semi-custom systems for training, post-training, reasoning, retrieval and serving.

不同的工作负载可能偏好不同的加速器,包括XPU、GPU、CPU和LPU。GPU系统可以与半定制系统协同工作,用于训练、后训练、推理、检索和服务。

NVLink Fusion addresses these challenges through a unified architecture. XPU- and GPU-based systems such as Vera Rubin NVL72 can share rack footprints, networking, cooling, power delivery and management systems. Operators can move forward with buildout while deferring the precise silicon mix, then reprovision capacity as workload demand, silicon supply and business priorities change.

NVLink Fusion通过统一架构应对这些挑战。基于XPU和GPU的系统,如Vera Rubin NVL72,可以共享机架占地面积、网络、冷却、供电和管理系统。运营商可以推进建设,同时推迟确定具体的硅片组合,然后根据工作负载需求、硅片供应和业务优先级的变化重新配置容量。

“The value of the NVLink Fusion program is … [customers] can deploy their rack-level solution with the NVIDIA GPU, and then they can decouple the development of their XPU and put it at a different pace,” said Vince Hu, corporate senior vice president and general manager of the data center and computing business group at MediaTek.

“NVLink Fusion计划的价值在于……客户可以部署他们的机架级解决方案,使用NVIDIA GPU,然后他们可以将XPU的开发解耦,并以不同的节奏进行,”联发科技企业高级副总裁兼数据中心和计算业务群总经理Vince Hu表示。

“NVLink Fusion allows the hyperscalers or the custom ASIC designers to integrate their own custom CPU or XPU and bridges the NVIDIA technology with a third-party process to create a unified rack-scale architecture,” said Lie-Szu Juang, chair and chief strategy officer at GUC.

“NVLink Fusion允许超大规模或定制ASIC设计者集成他们自己的定制CPU或XPU,并将NVIDIA技术与第三方流程桥接,创建统一的机架级架构,”GUC董事长兼首席战略官Lie-Szu Juang表示。

Designed, Validated and Operated as a Factory

设计、验证并作为工厂运营

Factory buildout is expensive, and mistakes can require costly rework. Infrastructure must be validated before construction begins. NVLink Fusion aligns with the NVIDIA DSX reference architecture for AI factories: codesigning buildings, power, cooling, compute and networking. The NVIDIA Omniverse DSX AI Factory Blueprint provides a digital twin and open reference design for gigawatt-scale AI factories, enabling partners to model facilities and technology together before deployment.

工厂建设成本高昂,错误可能导致昂贵的返工。基础设施必须在建设开始前进行验证。NVLink Fusion与NVIDIA DSX参考架构对齐,用于AI工厂:协同设计建筑、电力、冷却、计算和网络。NVIDIA Omniverse DSX AI工厂蓝图提供了千兆瓦级AI工厂的数字孪生和开放参考设计,使合作伙伴能够在部署前对设施和技术进行联合建模。

At the rack level, serviceability is part of performance. Reference compute trays feature 100% liquid cooling with no fans, cables or hoses, and allow trays to be removed while the rest of the rack remains operational. NVLink Switch trays are also liquid cooled and support continued operation during service.

在机架层面,可维护性是性能的一部分。参考计算托盘采用100%液冷,无风扇、电缆或软管,并允许在机架其余部分保持运行的情况下移除托盘。NVLink Switch托盘也采用液冷,并在维护期间支持持续运行。

“With NVLink Fusion we can use proven NVL72 rack design to have time-to-market, and we can have access to multiple suppliers to help us to deliver more into the hands of our customers,” said CC Lee, senior hardware development manager at Annapurna Labs, an Amazon company.

“借助NVLink Fusion,我们可以利用经过验证的NVL72机架设计来缩短上市时间,并且我们可以接触多家供应商,以帮助我们向客户交付更多产品,”亚马逊旗下Annapurna Labs高级硬件开发经理CC Lee表示。

Software completes the factory. NVIDIA NCCL for distributed workloads, NVIDIA Dynamo and NIXL for disaggregation and NVIDIA Mission Control for cluster management, telemetry and debugging help operators run mixed AI infrastructure as a coordinated system.

软件完善了工厂。NVIDIA NCCL 用于分布式工作负载,NVIDIA Dynamo 和 NIXL 用于分解,NVIDIA Mission Control 用于集群管理、遥测和调试,帮助运营商将混合 AI 基础设施作为协调系统运行。

With NVLink Fusion, XPUs can now meet a world-class AI platform, enabling hyperscalers and AI-native companies to build unified, semi-custom AI factories that combine the strengths of many builders into infrastructure no one company could build alone.

借助 NVLink Fusion,XPU 现在可以满足世界级 AI 平台的需求,使超大规模企业和 AI 原生公司能够构建统一的、半定制的 AI 工厂,将众多构建者的优势结合到任何单一公司都无法独自构建的基础设施中。

Learn more about NVLink Fusion.

了解更多关于 NVLink Fusion 的信息。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近