英伟达发布 Groq 3 LPX,Vera Rubin 推理性能提升 4 倍
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
做推理和智能体基础设施的同学必看,英伟达把长上下文 token 生成速度拉到每秒 3400 token,是竞品 4 倍,赶紧评估下对你的 agent 链路意味着什么。
The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.
AI推理的下一个时代不会由单一的突破性芯片、网络或系统来定义。它将由AI工厂每一层如何协同工作来定义。这就是为什么NVIDIA正在扩展Vera Rubin NVL72,为智能体系统提供快速令牌生成。
Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, it delivered 3,400 output tokens per second for 100,000-token long-context use cases critical to agentic systems, 4x faster than the nearest alternative platform.
今天宣布,NVIDIA Vera Rubin机架级系统NVIDIA Groq 3 LPX已全面投产。在运行开源智能体模型Gemma 4 31B的Artificial Analysis基准测试中,它为对智能体系统至关重要的10万令牌长上下文用例提供了每秒3400个输出令牌,比最接近的替代平台快4倍。
Industry partners worldwide are adopting Vera Rubin platform solutions. SpaceXAI announced that NVIDIA Vera CPUs will power its next generation of agentic AI. CoreWeave has deployed into production Spectrum-X Multiplane, which connects NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks. Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX.
全球行业合作伙伴正在采用Vera Rubin平台解决方案。SpaceXAI宣布,NVIDIA Vera CPU将为其下一代智能体AI提供动力。CoreWeave已部署生产级Spectrum-X Multiplane,该网络使用多个并行交换机连接NVIDIA Vera Rubin机架,提供高带宽、扁平且无损耗的AI网络。Nebius是首家采用NVIDIA Groq 3 LPX的AI云。
As AI shifts from training to reasoning and agentic, inference has become the new frontier. Agentic AI systems are generating more tokens, processing dramatically larger context windows and increasingly collaborating with other AI systems to solve complex problems.
随着AI从训练转向推理和智能体,推理已成为新的前沿。智能体AI系统正在生成更多令牌,处理更大的上下文窗口,并越来越多地与其他AI系统协作解决复杂问题。
These workloads demand a new class of infrastructure optimized not just for performance but for throughput, responsiveness and economics at unprecedented scale.
这些工作负载需要一类新的基础设施,不仅针对性能进行优化,还要在前所未有的规模上优化吞吐量、响应速度和经济性。
At the Hot Chips conference this week in Palo Alto, California, NVIDIA is showcasing how extreme codesign is reshaping the AI factory from end to end. By architecting compute, networking and inference acceleration as a unified system, NVIDIA is helping customers build infrastructure purpose-built for the emerging demands of long-context inference and multi-agent systems.
本周在加利福尼亚州帕洛阿尔托举行的Hot Chips会议上,NVIDIA正在展示极端协同设计如何端到端重塑AI工厂。通过将计算、网络和推理加速架构为统一系统,NVIDIA正在帮助客户构建专为长上下文推理和多智能体系统的新兴需求而设计的基础设施。
Extreme Codesign Optimizes for Performance
极端协同设计优化性能
Extreme codesign is the guiding principle behind NVIDIA platforms. Vera Rubin is engineered to accelerate inference as agents reason over increasingly long sequences.
极端协同设计是NVIDIA平台背后的指导原则。Vera Rubin旨在加速推理,因为智能体在越来越长的序列上进行推理。
NVIDIA Spectrum-X Ethernet moves those massive data flows efficiently across AI factories, and NVIDIA Groq 3 LPX is built to generate tokens at ultrafast speeds. Together, they show how NVIDIA is optimizing every stage of the AI pipeline, from context and communication to generation, as part of a single, integrated AI factory architecture.
NVIDIA Spectrum-X以太网在AI工厂中高效传输这些大规模数据流,而NVIDIA Groq 3 LPX旨在以超快速度生成令牌。它们共同展示了NVIDIA如何优化AI管道的每个阶段,从上下文和通信到生成,作为单一、集成的AI工厂架构的一部分。
NVIDIA Groq 3 LPX brings a new low-latency inference architecture designed to work alongside Vera Rubin NLV72, the most versatile AI factory platform, helping enterprises and cloud providers deliver the low latency, extreme throughput and scalable economics required for agentic applications.
NVIDIA Groq 3 LPX 引入了一种新的低延迟推理架构,旨在与 Vera Rubin NVL72(最通用的 AI 工厂平台)协同工作,帮助企业和云提供商实现代理应用所需的低延迟、极致吞吐量和可扩展的经济性。
Breakthrough performance comes not from optimizing individual components in isolation, but from codesigning every layer of the stack. From networking and context processing to large-scale inference, NVIDIA’s full-stack platform turns AI factories into integrated engines for intelligence, built to turn ever-growing volumes of tokens into revenue.
突破性的性能并非来自孤立地优化单个组件,而是来自对每一层技术栈的协同设计。从网络和上下文处理到大规模推理,NVIDIA 的全栈平台将 AI 工厂转变为智能引擎,旨在将不断增长的令牌量转化为收入。
Tuesday, Aug. 24, 8:00 a.m. PT
太平洋时间 8 月 24 日(周二)上午 8:00
NVIDIA Partners Adopt Vera Rubin for Lowest Token Costs
NVIDIA 合作伙伴采用 Vera Rubin 实现最低令牌成本
Nebius, a leading AI cloud, is first to adopt NVIDIA Groq 3 LPX, giving developers access to leading token generation speeds for highly responsive agentic AI applications.
领先的 AI 云服务商 Nebius 率先采用 NVIDIA Groq 3 LPX,为开发者提供领先的令牌生成速度,以支持高响应性的代理 AI 应用。
Adding NVIDIA Groq 3 LPX to NVIDIA Vera Rubin NVL72 in Nebius Token Factory will boost inference performance so developers can build highly interactive agents, coding systems and other real-time AI experiences at scale.
在 Nebius Token Factory 中将 NVIDIA Groq 3 LPX 添加到 NVIDIA Vera Rubin NVL72 将提升推理性能,使开发者能够大规模构建高度交互的代理、编码系统和其他实时 AI 体验。
Connecting NVIDIA Vera Rubin racks, CoreWeave is deploying Spectrum-X Multiplane in production, unlocking advances for its AI cloud infrastructure.
连接 NVIDIA Vera Rubin 机架,CoreWeave 正在生产环境中部署 Spectrum-X Multiplane,为其 AI 云基础设施带来进步。
Tuesday, Aug. 24, 8:00 a.m. PT
太平洋时间 8 月 24 日(周二)上午 8:00
SpaceXAI Adopts NVIDIA Vera CPUs for Agentic AI
SpaceXAI 采用 NVIDIA Vera CPU 用于代理 AI
SpaceXAI plans to build and scale its future AI architecture around NVIDIA Vera Rubin, from data centers on Earth to orbital satellites. The company plans to deploy NVIDIA Vera CPUs to accelerate the CPU-intensive work behind agentic AI, including orchestration, tool use, code execution, data processing and simulation.
SpaceXAI 计划围绕 NVIDIA Vera Rubin 构建和扩展其未来的 AI 架构,从地球上的数据中心到轨道卫星。该公司计划部署 NVIDIA Vera CPU 来加速代理 AI 背后的 CPU 密集型工作,包括编排、工具使用、代码执行、数据处理和模拟。
The SpaceXAI partnership extends NVIDIA’s full-stack AI platform to SpaceXAI, bringing together Vera CPUs, NVIDIA accelerated computing, networking and software to advance AI at unprecedented scale.
SpaceXAI 合作将 NVIDIA 的全栈 AI 平台扩展到 SpaceXAI,整合 Vera CPU、NVIDIA 加速计算、网络和软件,以前所未有的规模推进 AI。
Designed for the agentic era, Vera Rubin provides leading per-core performance, exceptional memory bandwidth and predictable performance under load, helping agents complete tasks faster and keeping valuable GPU infrastructure fully utilized.
专为代理时代设计,Vera Rubin 提供领先的单核性能、卓越的内存带宽和负载下的可预测性能,帮助代理更快完成任务,并保持宝贵的 GPU 基础设施得到充分利用。
Tuesday, Aug. 24, 8:00 a.m. PT
太平洋时间 8 月 24 日(周二)上午 8:00
NVIDIA Groq 3 LPX: The Interactive AI Inference Accelerator
NVIDIA Groq 3 LPX:交互式 AI 推理加速器
Codesigned with the Vera Rubin NVL72 platform, NVIDIA Groq 3 LPX is helping AI factories deliver tokens at the lowest latency for agentic workloads.
与 Vera Rubin NVL72 平台协同设计,NVIDIA Groq 3 LPX 帮助 AI 工厂以最低延迟为代理工作负载提供令牌。
Agentic AI is creating a new performance challenge: decode latency. As AI agents reason, use tools and interact with other systems, they generate responses one token at a time, causing even tiny delays to multiply across complex chains of work. To keep agents operating at the pace users expect, NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with specialized acceleration for token generation.
代理型人工智能正带来一项新的性能挑战:解码延迟。由于AI代理在推理、使用工具及与其他系统交互时,会逐词生成响应,即使是微小的延迟也会在复杂的工作链中成倍放大。为使代理以用户期望的速度运行,NVIDIA Groq 3 LPX扩展了Vera Rubin NVL72平台,专门加速令牌生成。
NVIDIA Rubin GPUs handle large-scale context processing while LPX accelerates latency-sensitive decode workloads. The result is faster, more predictable token generation that helps AI factories deliver responsive reasoning, smoother agent interactions and greater infrastructure efficiency.
NVIDIA Rubin GPU负责处理大规模上下文,而LPX则加速对延迟敏感的解码工作负载。其结果是更快、更可预测的令牌生成,帮助AI工厂实现响应式推理、更流畅的代理交互以及更高的基础设施效率。
Together, Rubin GPUs and LPUs are designed to eliminate the traditional tradeoff between speed and throughput, helping AI providers deliver responsive, large-scale inference for the next generation of agentic AI applications.
Rubin GPU与LPU共同设计,旨在消除速度与吞吐量之间的传统权衡,帮助AI服务提供商为下一代代理型AI应用提供响应迅速的大规模推理。
Building the Token Factory
构建令牌工厂
As the industry shifts from model training to serving intelligence at scale, infrastructure must evolve into what NVIDIA describes as a “token factory” capable of delivering performance, throughput, intelligence integrity and economic efficiency simultaneously. Agentic AI systems increasingly communicate with other AI systems, access multiple data sources and maintain large amounts of context, creating unprecedented demand for fast inference.
随着行业从模型训练转向大规模智能服务,基础设施必须演进为NVIDIA所称的“令牌工厂”,能够同时实现性能、吞吐量、智能完整性和经济效率。代理型AI系统日益与其他AI系统通信、访问多个数据源并维护大量上下文,对快速推理产生了前所未有的需求。
NVIDIA Groq 3 LPX was designed for exactly these workloads. As an extension of the Vera Rubin NVL72, it enables ultrafast responsiveness even across massive context windows while helping service providers maximize throughput and infrastructure utilization.
NVIDIA Groq 3 LPX正是为这些工作负载而设计。作为Vera Rubin NVL72的扩展,它即使在巨大的上下文窗口下也能实现超快响应,同时帮助服务提供商最大化吞吐量和基础设施利用率。
Extreme Codesign for Inference
面向推理的极致协同设计
Unlike standalone accelerators, NVIDIA Groq 3 LPX combines the strengths of GPUs and LPUs through extreme codesign. Rubin GPUs and LPUs jointly compute every layer of an AI model, enabling new levels of inference performance for agentic workloads.
与独立加速器不同,NVIDIA Groq 3 LPX通过极致协同设计结合了GPU和LPU的优势。Rubin GPU与LPU共同计算AI模型的每一层,为代理型工作负载带来新的推理性能水平。
At scale, fleets of LPUs operate as a giant processor optimized for deterministic inference. A rack-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, creating a highly efficient inference engine built for modern AI factories.
在规模上,LPU集群作为一个为确定性推理优化的大型处理器运行。机架级NVIDIA Groq 3 LPX部署可包含256个LP30加速器,通过直接芯片间链路连接,构建出专为现代AI工厂设计的高效推理引擎。
Designed for the Agentic AI Era
专为代理型AI时代设计
As reasoning models grow and agentic workflows generate ever more tokens, the infrastructure required to serve them must evolve. NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with a purpose-built inference architecture designed to maximize responsiveness, throughput and efficiency, helping power the next generation of AI factories.
随着推理模型的增长和代理工作流生成越来越多的令牌,为其提供服务的底层基础设施必须不断演进。NVIDIA Groq 3 LPX 扩展了 Vera Rubin NVL72 平台,采用专为推理设计的架构,旨在最大化响应速度、吞吐量和效率,助力驱动下一代 AI 工厂。
And this is only the beginning, more optimizations, more models, more performance when paired with Vera Rubin NVL72 — new levels of throughput and interactivity are coming. Stay tuned.
而这仅仅是开始,与 Vera Rubin NVL72 搭配使用时,将带来更多优化、更多模型和更高性能——新的吞吐量和交互性水平即将到来。敬请期待。
Tuesday, Aug. 24, 8:00 a.m. PT
太平洋时间 8 月 24 日(周二)上午 8:00
NVIDIA Spectrum-X Multiplane Enables Massive AI Factory Scale on a Flatter, More Resilient Network
NVIDIA Spectrum-X Multiplane 在更扁平、更具弹性的网络上实现大规模 AI 工厂扩展
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力