Nvidia Groq 3 LPX 系统小模型解码速度超3400 tok/s
Nvidia's Groq 3 LPX system excels at fast decode of small models (>3,400 tok/s),…
Nvidia's Groq 3 LPX system excels at fast decode of small models (>3,400 tok/s), according to a @ArtificialAnlys benchmarks of Gemma 4 31B. LPX does this by using a small amount (128 GB) of ultrafast SRAM in place of HBM.
Nvidia proposes combining LPUs with GPUs to bring this speed to frontier-scale models. A look at how that (might) work below.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力