跳到主内容
@wquguru
精选75Chubby♨️行业动态

Cerebras发布CS-4,推理性能翻倍

This is seriously impressive: Cerebras just unveiled CS-4, and nearly doubled it…

原文
发到 X

This is seriously impressive: Cerebras just unveiled CS-4, and nearly doubled its AI inference performance without moving to a new process node.

这确实令人印象深刻:Cerebras 刚刚发布了 CS-4,在不更换新工艺节点的情况下,几乎将 AI 推理性能翻了一番。

Same gigantic 5nm wafer. Same 4 trillion transistors. Same 900,000 AI cores. Instead, Cerebras redesigned the power delivery and cooling, allowing the wafer to run at twice the clock speed. The result per WSE-3 Turbo:

同样的巨型 5nm 晶圆。 同样的 4 万亿晶体管。 同样的 90 万 AI 核心。 相反,Cerebras 重新设计了供电和冷却系统,使晶圆能够以两倍的时钟速度运行。 每片 WSE-3 Turbo 的结果:

  • 250 PFLOPs of AI compute - 43.2 PB/s of memory bandwidth - 2.4 Tb/s of I/O bandwidth
  • 250 PFLOPs 的 AI 计算能力 - 43.2 PB/s 的内存带宽 - 2.4 Tb/s 的 I/O 带宽

A single CS-4 rack combines three wafers for 750 PFLOPs and 129.6 PB/s of memory bandwidth.

单个 CS-4 机架结合三片晶圆,提供 750 PFLOPs 的计算能力和 129.6 PB/s 的内存带宽。

On GPT-OSS-120B, Cerebras reports more than 4,400 tokens per second per user, and up to 30x faster inference than GPU-based systems.

在 GPT-OSS-120B 上,Cerebras 报告每位用户每秒超过 4,400 个令牌,推理速度比基于 GPU 的系统快达 30 倍。

So yeah, intelligence not only too cheap to meter but also too fast to keep up with

所以,是的,智能不仅便宜到无需计量,而且快到难以跟上。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近