跳到主内容
@wquguru
精选65r/LocalLLaMA(Reddit)技巧与观点

利用 Strata 与 Qwen3.8 Next 优化低配硬件推理性能

Anyone sitting on a lot of slow system memory and a modest GPU.. try Strata + Qwen3.8 Next.

原文
发到 X

IQ3_XXS weights are just under 80GB and my slowww DDR4+7900XTX is stabilizing around 45-70/s (sometimes higher while coding depending on mtp). Looking online I'm seeing similar results for users with 12GB and 16GB cards, and significantly faster numbers for owners of DDR5.

IQ3_XXS 权重文件不到 80GB,我那慢吞吞的 DDR4+7900XTX 平台稳定在 45-70 token/s(有时更高,取决于 mtp)。我在网上看到拥有 12GB 和 16GB 显存卡的用户也有类似结果,而使用 DDR5 的用户速度则显著更快。

(In comparison, Llama CPP with tuning was maxing out around 22.5t/s on the same rig. Quality seems reliably superior (I wouldn't recommend the Q2 weights though))

(相比之下,经过调优的 Llama CPP 在同一套配置上最高只能达到约 22.5t/s。质量似乎确实更可靠(不过我不推荐 Q2 权重))

Seriously. Ask <LLM of your choosing> to set it up for your specs. If 27B doesnt fit well for you, here's a shot at beating it.

说真的。让 <你选择的 LLM> 根据你的硬件规格进行设置。如果 27B 模型不适合你,这里有一个超越它的方案。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件