精选70MiniMax (official)产品发布/更新
Fireworks AI 在 B200 上实现稀疏注意力 980 TFLOP/s
Sparse attention only matters if the systems can preserve the theoretical gains.
Sparse attention only matters if the systems can preserve the theoretical gains.
@FireworksAI_HQ new M3 kernel on Blackwell uses a KV-stationary design to read each selected block once, reaching ~980 TFLOP/s on a B200.
Read the full breakdown below for a closer look at the kernel design and optimizations. 👇
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力