精选75Ollama(GitHub Releases)AI 编程与模型
Ollama v0.32.6:Qwen3.5 在 Apple GPU 提速,OpenAI 流式格式修正
v0.32.6
What's Changed
- Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically
- /v1/chat/completions streaming now matches OpenAI's wire format: role only on the first chunk, finish_reason on its own chunk,
- and usage in a separate chunk with stream_options.include_usage.
- Truncated OpenAI responses now report finish_reason: "length" instead of "tool_calls".
- ollama run kimi-k3 now offers kimi-k3:cloud for cloud-only models that publish no default tag, instead of failing.
- TUI fixes: pipe-delimited prose no longer renders as a table, Enter accepts the highlighted @ file completion, and /prompt
- scrolling is no longer laggy.
- Experimental image generation has been temporarily removed. Continue using 0.32.5 for image generation support
- Updated the MLX and llama.cpp engines.
Full Changelog: v0.32.5...v0.32.6-rc0
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力