跳到主内容
@wquguru
精选90r/LocalLLaMA(Reddit)模型发布/更新

Qwen3.8-Flash-Next发布GSQ-RCO量化版及专家剪枝Code

[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw

原文
发到 X
推荐理由

本地部署者福音,用极低显存实现了接近全精度性能的代码能力,实测数据详实,值得收藏压测。

We have released Qwen3.8-Flash-Next quantized with GSQ and RCO, together with a second, capability-targeted build in which half of the model's experts have been removed.

我们发布了采用 GSQ 和 RCO 量化的 Qwen3.8-Flash-Next,以及一个以能力为目标的第二版构建,其中移除了模型一半的专家。

Flash-Next is a sparse mixture-of-experts model: 512 routed experts per layer across 48 layers, 176.9B parameters, 354 GB at BF16.

Flash-Next 是一个稀疏混合专家(MoE)模型:48 层中每层有 512 个路由专家,共 1769 亿参数,BF16 精度下占用 354 GB。

What's inside

内部结构

  • Four quantized GGUFs, 2.40 to 3.50 bpw (66.4 to 83.6 GB), and the BF16 vision projector
  • Expert-pruned Coder GGUF, 58.4 GB in total, of which 29.6 GB must remain resident
  • GSQ (Gumbel-Softmax Quantization): post-training scalar quantization that jointly learns grid assignments and group scales, closing most of the gap between scalar and vector quantization at low bit-widths while remaining deployable in standard GGUF types
  • RCO (Riemannian Constrained Optimization): enforces exact budgets by gradient descent on the task loss, without per-constraint tuning. It serves two roles in this release: assigning a quantization type to every tensor, and selecting which experts to retain in the Coder build, where it enforces several exact budgets simultaneously, one per layer
  • 四个量化的 GGUF 文件,位宽在 2.40 到 3.50 bpw 之间(66.4 至 83.6 GB),以及 BF16 精度的视觉投影器
  • 经过专家剪枝的 Coder GGUF,总大小 58.4 GB,其中必须有 29.6 GB 常驻内存
  • GSQ(Gumbel-Softmax 量化):一种后训练标量量化方法,联合学习网格分配和组缩放比例,在低位宽下缩小了标量量化与向量量化之间的差距,同时仍可在标准 GGUF 类型中部署
  • RCO(黎曼约束优化):通过对任务损失进行梯度下降来强制执行精确预算,无需针对每个约束进行调整。它在本次发布中承担两个角色:为每个张量分配量化类型,以及在 Coder 构建中选择保留哪些专家,在此过程中它同时强制执行多个精确预算,每层一个

Results:

结果:

At 3.50 bpw the model matches the BF16 base on every benchmark evaluated.

在 3.50 bpw 时,该模型在所有评估基准上均与 BF16 基础模型持平。

  • IQ3_S (3.50 bpw, 83.6 GB): AIME25 100.00, GPQA-Diamond 92.93 against 91.92 for BF16, LiveCodeBench v6 86.86 against 87.43. Task average 93.26 against 93.12.
  • IQ3_XXS (3.00 bpw, 75.8 GB): AIME25 100.00, GPQA-Diamond 91.41, LiveCodeBench v6 86.29
  • Q2_0 (2.40 bpw, 66.4 GB): zero-shot average 78.00, above the BF16 value of 76.94, at approximately one fifth of the size
  • IQ3_S(3.50 bpw,83.6 GB):AIME25 得分 100.00,GPQA-Diamond 得分 92.93(对比 BF16 的 91.92),LiveCodeBench v6 得分 86.86(对比 87.43)。任务平均分 93.26(对比 93.12)。
  • IQ3_XXS(3.00 bpw,75.8 GB):AIME25 得分 100.00,GPQA-Diamond 得分 91.41,LiveCodeBench v6 得分 86.29
  • Q2_0(2.40 bpw,66.4 GB):零样本平均分 78.00,高于 BF16 的 76.94,且大小仅为原来的约五分之一

Coder (capability pruned model):

Coder(能力剪枝模型):

Instead of storing every parameter at lower precision, half of the routed experts are removed from the model: 256 of 512 per layer, selected by RCO optimising the KL divergence against the unpruned model. The retained weights remain at 3.5 bpw. Pruning and quantization compound, and the combined effect is an average of 1.89 bits per parameter of the original transformer. The averaged bitwidth amortises the removed experts over the original parameter count, and therefore expresses the joint effect of pruning and quantization. No individual weight is stored at 1.89 bits.

与其以较低精度存储所有参数,不如从模型中移除一半的路由专家:每层 512 个中的 256 个,通过 RCO 优化与未剪枝模型的 KL 散度来选择。保留的权重仍保持在 3.5 bpw。剪枝和量化效应叠加,综合效果使原始 Transformer 每个参数的平均位宽降至 1.89 位。平均位宽将移除的专家分摊到原始参数数量上,从而表达了剪枝和量化的联合效应。没有任何单个权重以 1.89 位存储。

The practical consequence is that a 176.9B-parameter model has a resident working set of 29.6 GB, since the n-gram shard is a lookup table and may be served from disk. This is within the capacity of a single 32 GB accelerator.

实际结果是,一个拥有 1769 亿参数的模型,其常驻工作集大小为 29.6 GB,因为 n-gram 分片是查找表,可以从磁盘提供服务。这在一个 32 GB 加速器的容量范围内。

  • SWE-bench Verified: 75.60 against 82.80 for BF16, retaining 91.3%
  • LiveCodeBench v6: 86.28 against 87.43, retaining 98.7%
  • SWE-bench Verified:得分 75.60(对比 BF16 的 82.80),保留了 91.3%
  • LiveCodeBench v6:得分 86.28(对比 87.43),保留了 98.7%

Both measured at xhigh reasoning effort.

均在 xhigh 推理努力下测量。

Links

链接

  • Quantized models: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
  • Coder (expert-pruned): https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
  • GSQ: paper https://arxiv.org/abs/2604.18556 | code https://github.com/IST-DASLab/GSQ
  • RCO: paper https://arxiv.org/abs/2605.00649 | code https://github.com/IST-DASLab/RCO
  • 量化模型:https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
  • Coder(专家剪枝版):https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
  • GSQ:论文 https://arxiv.org/abs/2604.18556 | 代码 https://github.com/IST-DASLab/GSQ
  • RCO:论文 https://arxiv.org/abs/2605.00649 | 代码 https://github.com/IST-DASLab/RCO

Both repositories ship the complete per-tensor RCO allocation.

两个仓库均提供完整的逐张量 RCO 分配方案。

The Coder build is an experimental release and feedback is welcome, particularly on capabilities that were not represented in the calibration mixture. Requests for models to quantize or prune are also welcome.

Coder 构建版本为实验性发布,欢迎反馈,特别是针对校准混合数据中未体现的能力方面的反馈。也欢迎提出需要量化或剪枝的模型请求。

From the ISTA Deep Algorithms and Systems Lab.

来自 ISTA 深度算法与系统实验室。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件