利用 Strata 与 Qwen3.8 Next 优化低配硬件推理性能
Anyone sitting on a lot of slow system memory and a modest GPU.. try Strata + Qwen3.8 Next.
IQ3_XXS weights are just under 80GB and my slowww DDR4+7900XTX is stabilizing around 45-70/s (sometimes higher while coding depending on mtp). Looking online I'm seeing similar results for users with 12GB and 16GB cards, and significantly faster numbers for owners of DDR5.
IQ3_XXS 权重文件不到 80GB,我那慢吞吞的 DDR4+7900XTX 平台稳定在 45-70 token/s(有时更高,取决于 mtp)。我在网上看到拥有 12GB 和 16GB 显存卡的用户也有类似结果,而使用 DDR5 的用户速度则显著更快。
(In comparison, Llama CPP with tuning was maxing out around 22.5t/s on the same rig. Quality seems reliably superior (I wouldn't recommend the Q2 weights though))
(相比之下,经过调优的 Llama CPP 在同一套配置上最高只能达到约 22.5t/s。质量似乎确实更可靠(不过我不推荐 Q2 权重))
Seriously. Ask <LLM of your choosing> to set it up for your specs. If 27B doesnt fit well for you, here's a shot at beating it.
说真的。让 <你选择的 LLM> 根据你的硬件规格进行设置。如果 27B 模型不适合你,这里有一个超越它的方案。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力