跳到主内容
@wquguru
精选72Ollama(GitHub Releases)AI 编程与模型

Ollama v0.32.10-rc1:repeat_penalty 默认 1.0,NVFP4 提速

v0.32.10

原文
发到 X

What's Changed

变更内容

  • Models that don't set a repeat_penalty now default to 1.0 (off) instead of 1.1, matching other engines and speeding up speculative decoding; set a per-model parameter if an older model repeats itself.
  • Faster prefill on NVFP4 MLX models with a global scale, about 7–8% on Qwen3.6 and Muse Glimmer.
  • Fixed blob verification being skipped when an OCI manifest's config and layer share a digest.
  • 未设置 repeat_penalty 的模型现在默认值为 1.0(关闭),而不是 1.1,与其他引擎保持一致,并加速推测解码;如果旧模型出现重复,请设置每模型参数。
  • 在具有全局尺度的 NVFP4 MLX 模型上加快预填充,Qwen3.6 和 Muse Glimmer 上约提升 7-8%。
  • 修复了当 OCI 清单的配置和层共享摘要时跳过 blob 验证的问题。

New Contributors

新贡献者

  • @vigneshakaviki made their first contribution in #15504
  • @vigneshakaviki 在 #15504 中做出了他们的首个贡献

Full Changelog: v0.32.8...v0.32.10-rc1

完整变更日志:v0.32.8...v0.32.10-rc1

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近