跳到主内容
@wquguru
精选75Hacker News Best(web_list)行业动态

Cerebras 发布 CS-4 推理系统,速度较 GPU 快 30 倍

Cerebras CS-4

原文
发到 X

The Fastest AI

最快的AI

Just Got Faster.

如今更快了。

Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to deploy hyperscale capacity. It is the architecture for frontier AI.​

隆重推出全新Cerebras CS-4,这是一款革命性的机架级解决方案,与GPU相比,推理速度最高可提升30倍,经济性更佳,并提供了部署超大规模容量的简单路径。它是面向前沿AI的架构。

Three WSE-3 Turbo per System​

每系统配备三个WSE-3 Turbo

Each wafer delivers up to 2x the speed of the previous generation​

每个晶圆的速度比上一代最高提升2倍

More Performance per Wafer​

每个晶圆性能更高

All new power, cooling, and I/O unleashes even more performance per wafer​

全新的电源、冷却和I/O设计释放了每个晶圆更多的性能

Nexus Rack-Scale Platform

Nexus机架级平台

Enables rapid deployment in hyperscale datacenters​

支持在超大规模数据中心快速部署

Up to 30x faster than GPUs​

比GPU快高达30倍

Powered by WSE-Turbo, CS-4 delivers up to 30x faster inference compared to GPU systems, setting a new record for the fastest inference available in production.​

由WSE-Turbo驱动,CS-4与GPU系统相比,推理速度最高可提升30倍,创下了生产环境中可用最快推理的新纪录。

Higher ultrafast throughput

更高的超快吞吐量

The CS-4 solution shifts the inference Pareto frontier, delivering up to 10x more throughput per watt than CS-3 while generating tokens up to 30x faster than production GPU systems. The result is a system designed to deliver both throughput and interactivity.​

CS-4解决方案推动了推理帕累托前沿的转移,每瓦吞吐量比CS-3提升高达10倍,同时生成令牌的速度比生产级GPU系统快30倍。其结果是设计用于同时提供吞吐量和交互性的系统。

Frontier-ready architecture

面向前沿的架构

By reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters, preserving interactive decode performance at unprecedented scale.​

通过将晶圆间互连延迟降至2微秒,CS-4在超过10万亿参数的模型上每秒可生成超过1000个令牌,在前所未有的规模下保持了交互式解码性能。

BUILT FOR HYPERSCALE​

为超大规模而生

CS-4 is the first iteration of the new Cerebras Nexus Platform Architecture. It is built around a modular concept with three foundational elements: Compute, Power, and I/O – each with significant innovation to simplify manufacturing, deployment, maintenance, and upgrades.​

CS-4是全新Cerebras Nexus平台架构的首个迭代。它围绕模块化概念构建,包含三个基础元素:计算、电源和I/O——每个元素都有重大创新,以简化制造、部署、维护和升级。

Modular compute backpack design

模块化计算背包设计

Cerebras has fundamentally re-imagined the server. Each Wafer-Scale Backpack is a self-contained assembly thatfolds the wafer, power conversion, direct liquid cooling, high-speed I/O, and control electronics into a compact 3D package with 50% fewer components. This design simplifies manufacturing and reduces deployment time from days to hours.​

Cerebras从根本上重新构想了服务器。每个晶圆级背包是一个自包含的组件,将晶圆、电源转换、直接液体冷却、高速I/O和控制电子设备折叠成一个紧凑的3D封装,组件数量减少50%。这种设计简化了制造,并将部署时间从数天缩短至数小时。

High-density power delivery

高密度电源传输

With power delivery just 0.5 millimeters away from the processor - roughly 100x closer than the roughly 50mm of conventional GPU boards - CS-4 nearly eliminates board-level power loss. This enables the delivery of twice as much power to the WSE-3T, enabling higher operating frequencies and faster token generation.​

电源传输距离处理器仅0.5毫米——比传统GPU板的约50毫米近了约100倍——CS-4几乎消除了板级功率损耗。这使得能够向WSE-3T提供两倍的功率,从而实现更高的工作频率和更快的令牌生成。

Next-gen wafer I/O interface

下一代晶圆I/O接口

CS-4 introduces a new programmable I/O subsystem that doubles I/O bandwidth and reduces latency, benefitting both aggregated and disaggregated solutions. The Wafer I/O Module also enables wafers to be linked within and across racks without a switch, for wafer-to-wafer latency as low as two microseconds that is key to interactivity for models with tens of trillions of parameters.​

CS-4 引入了新的可编程 I/O 子系统,使 I/O 带宽翻倍并降低延迟,惠及聚合与解聚解决方案。晶圆 I/O 模块还使晶圆能够在机架内和跨机架间无需交换机即可连接,实现低至两微秒的晶圆间延迟,这对于拥有数十万亿参数模型的交互性至关重要。

Deploy infrastructure then compute

先部署基础设施,再部署计算

CS-4 separates the stable power, cooling, and network layer from its modular wafer-scale compute. The Cerebras PowerRack can be installed and facility-qualified before compute arrives. Compute backpacks then slide into place and connect to power, cooling, and data—reducing deployment from days to hours while simplifying service and future upgrades at hyperscale.​

CS-4 将稳定的电源、冷却和网络层与其模块化晶圆级计算分离。Cerebras PowerRack 可在计算设备到达前安装并通过设施验证。随后计算背包滑入到位,并连接电源、冷却和数据——将部署时间从数天缩短至数小时,同时简化了超大规模环境下的维护和未来升级。

CS-4 by the numbers

CS-4 数据一览

First CS-4 shipments begin this quarter.​

首批 CS-4 将于本季度开始发货。

Bring the fastest AI to your data center.​

将最快的 AI 引入您的数据中心。

Get startedDatasheet

开始使用 数据手册

FAQ

常见问题解答

What is Cerebras CS-4?

什么是 Cerebras CS-4?

How is CS4 different from CS-3?

CS-4 与 CS-3 有何不同?

How fast is CS-4?

CS-4 有多快?

How much throughput does CS-4 provide?

CS-4 提供多少吞吐量?

Why is CS-4 well suited for agentic AI?

为什么 CS-4 非常适合代理式 AI?

What is a Wafer-Scale Backpack?

什么是晶圆级背包?

How does the Nexus Platform Architecture simplify hyperscale deployment?

Nexus 平台架构如何简化超大规模部署?

What models and inference architectures does CS-4 support?

CS-4 支持哪些模型和推理架构?

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近