Qwen3.8-Next在6卡V100部署实测:引擎选型与MTP加速
Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3)
本地部署Qwen3.8-Next的多卡踩坑实录,特别是引擎避坑指南和MTP带来的显著提速效果,做私有化部署的同学可以直接参考其参数配置。
Hey everyone,
大家好,
After wrestling with hardware and engine issues for days, I finally got Qwen 3.8 Next running properly on my multi-GPU rig. Thought I’d share the setup journey, benchmarks, and thermal results for anyone trying something similar.
在与硬件和引擎问题搏斗了几天后,我终于让我的多 GPU 机器上成功运行了 Qwen 3.8 Next。我想分享一下设置过程、基准测试结果以及热表现数据,供有类似尝试的朋友参考。
Seeing all the ongoing memes on Reddit about multi-GPU setups turning into absolute space heaters and catching fire, I decided to run some rigorous thermal tests to see for myself.
看到 Reddit 上关于多 GPU 设置变成巨型暖风机甚至起火的梗图满天飞,我决定进行一些严格的热测试,亲自验证一下。
he Troubleshooting Odyssey
故障排除之旅
- PCIe Link Speed Issue: Right after installation, one of the cards dropped to PCIe Gen 1 x16. Spent about 8 hours over two days diagnosing and fixing it.
- Finding the Right Engine:
- Started with sglang-v100, but kept hitting continuous OOM crashes.
- Someone on Reddit previously suggested the pxa engine, but that threw errors as well.
- Eventually tried 1cat-vllm, spent some time tweaking it, and finally hit a stable run!
- Configuration:
- Running with TP2 PP3.
- Currently, speculative decoding is limited to speculative=1. Setting it to 2 throws an OOM due to memory constraints (might look into optimizing this later, but for now, it works).
- PCIe 链路速度问题:安装完成后不久,其中一张卡掉到了 PCIe Gen 1 x16。我花了两天时间,大约 8 小时来诊断并修复这个问题。
- 寻找合适的引擎:
- 起初使用 sglang-v100,但不断遇到持续的 OOM(内存溢出)崩溃。
- Reddit 上有人之前推荐过 pxa 引擎,但它也报错。
- 最终尝试了 1cat-vllm,经过一番调整,终于实现了稳定运行!
- 配置:
- 采用 TP2 PP3 配置。
- 目前,推测解码(speculative decoding)限制为 speculative=1。将其设置为 2 会因内存限制导致 OOM(稍后可能会研究优化这一点,但目前这样就能用)。
Context & Memory Stats
上下文与内存统计
Plaintext
纯文本
INFO: Available KV cache memory: 8.78 GiB INFO: GPU KV cache size: 531,288 tokens INFO: Maximum concurrency for 262,144 tokens per request: 2.03xINFO: Available KV cache memory: 8.78 GiB INFO: GPU KV cache size: 531,288 tokens INFO: Maximum concurrency for 262,144 tokens per request: 2.03xPerformance Benchmarks
性能基准测试
1. Prompt Processing (Prefill)
1. 提示词处理(预填充 Prefill)
| Input Length (Tokens) | Speed (tok/s) |
|---|---|
| 1,024 (1K) | 1,389 |
| 2,048 (2K) | 2,536 |
| 4,096 (4K) | 3,210 |
| 8,192 (8K) | 4,336 |
| 16,384 (16K) | 4,679 |
| 32,768 (32K) | 4,470 |
| 65,536 (64K) | 3,820 |
| 131,072 (131K) | 2,759 |
| 输入长度(Token) | 速度 (tok/s) |
|---|---|
| 1,024 (1K) | 1,389 |
| 2,048 (2K) | 2,536 |
| 4,096 (4K) | 3,210 |
| 8,192 (8K) | 4,336 |
| 16,384 (16K) | 4,679 |
| 32,768 (32K) | 4,470 |
| 65,536 (64K) | 3,820 |
| 131,072 (131K) | 2,759 |
2. Text Generation (MTP Comparison)
2. 文本生成(MTP 对比)
| Output Length (Tokens) | Base Speed (No MTP, tok/s) | Optimized Speed (MTP Enabled, tok/s) |
|---|---|---|
| 128 | 22.23 | 41.34 |
| 256 | 22.32 | 42.48 |
| 512 | 22.70 | 43.04 |
| 1024 | 22.86 | 43.28 |
| 2048 | 22.83 | 43.38 |
| Average | 22.59 | 42.70 |
| 输出长度(Token) | 基础速度(无 MTP, tok/s) | 优化速度(启用 MTP, tok/s) |
|---|---|---|
| 128 | 22.23 | 41.34 |
| 256 | 22.32 | 42.48 |
| 512 | 22.70 | 43.04 |
| 1024 | 22.86 | 43.28 |
| 2048 | 22.83 | 43.38 |
| 平均值 | 22.59 | 42.70 |
MTP nearly doubles generation throughput across the board.
MTP 几乎使整体生成吞吐量翻了一番。
Thermals & Acoustics
热表现与噪音
People often meme about multi-GPU rigs turning into space heaters or jet engines, so I ran a thorough thermal/stress test:
人们常调侃多 GPU 机器会变成暖风机或喷气发动机,所以我进行了一次彻底的热/压力测试:
- Stress Test: Ran gpu-burn continuously for 20 minutes.
- Thermal Equilibrium: Temperatures peaked at 65°C and stabilized right around 64°C.
- Fan Curve: Based on my fan control script, the fans were only running at around 76% at 64°C. The cards stay well under 65°C without even needing full blast.
- 压力测试:连续运行 gpu-burn 20 分钟。
- 热平衡:温度峰值为 65°C,随后稳定在 64°C 左右。
- 风扇曲线:根据我的风扇控制脚本,在64°C时风扇转速仅在76%左右。即使不需要全速运转,显卡温度也能保持在65°C以下。
Pretty happy with how stable, cool, and quiet this system turned out.
对这台系统最终呈现出的稳定、低温和静音表现非常满意。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力