跳到主内容
@wquguru
精选70Epoch AI行业动态多源精选 ×5

Nvidia Groq 3 LPX 系统小模型解码速度超3400 tok/s

Nvidia's Groq 3 LPX system excels at fast decode of small models (>3,400 tok/s),…

原文
发到 X

Nvidia's Groq 3 LPX system excels at fast decode of small models (>3,400 tok/s), according to a @ArtificialAnlys benchmarks of Gemma 4 31B. LPX does this by using a small amount (128 GB) of ultrafast SRAM in place of HBM.

Nvidia proposes combining LPUs with GPUs to bring this speed to frontier-scale models. A look at how that (might) work below.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →