跳到主内容
@wquguru
精选85elvis论文研究

NVIDIA 压缩 MoE 模型 Puzzle-75B,吞吐量翻倍

Banger compression paper from NVIDIA.

原文
发到 X

Banger compression paper from NVIDIA.

(bookmark it)

Bigger MoE models keep winning on quality, but serving them at interactive latency is still hard.

NVIDIA compresses the hybrid MoE Nemotron-3-Super into Puzzle-75B-A9B and roughly doubles interactive server throughput while holding quality.

Pay attention to the joint structural search. Heterogeneous MoE pruning, active-parameter budget, and Mamba pruning get optimized together rather than one at a time, wrapped in an iterative pipeline with distillation, RL, quantization, and a Multi-Token Prediction head.

Why does it matter?

On a single 8xB200 node it hits about 2x the parent's server throughput at matched user-throughput, and 1M-token concurrency on a single H100 climbs from 1 request to 8. Accuracy holds across reasoning, coding, long-context, and agentic benchmarks.

Cheaper serving with agentic capability intact changes what you can afford to run with these models.

Paper: https://arxiv.org/abs/2607.04371

Learn to build effective AI agents in our academy: https://academy.dair.ai/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近