跳到主内容
@wquguru
精选70MiniMax (official)产品发布/更新

Fireworks AI 在 B200 上实现稀疏注意力 980 TFLOP/s

Sparse attention only matters if the systems can preserve the theoretical gains.

原文
发到 X

Sparse attention only matters if the systems can preserve the theoretical gains.

@FireworksAI_HQ new M3 kernel on Blackwell uses a KV-stationary design to read each selected block once, reaching ~980 TFLOP/s on a B200.

Read the full breakdown below for a closer look at the kernel design and optimizations. 👇

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近