PrismML发布Bonsai 2 27B:98.2%性能保留的三元权重模型
Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
低比特压缩技术在保持旗舰级多模态能力上的突破,对追求端侧部署效率的工程团队极具参考价值。
LAUNCH
发布
1001011100 11010 1 001
1001011100 11010 1 001
Back to all posts
返回所有帖子
Introducing Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
介绍 Bonsai 2 27B:在 9 倍更小的体积中实现近乎无损的压缩
September 17, 2026
2026 年 9 月 17 日
•
•
PrismML
PrismML
Two months ago, we released our first Bonsai 27B models and showed that a 27B-class multimodal model could be compressed enough to run efficiently on a local device. Today, we’re releasing Ternary Bonsai 2 27B, our most capable model yet.
两个月前,我们发布了首批 Bonsai 27B 模型,并证明了一个 27B 级别的 multimodal(多模态)模型可以被压缩到足以在本地设备上高效运行的程度。今天,我们发布 Ternary Bonsai 2 27B,这是我们迄今为止最强大的模型。
Based on Qwen3.8 27B, Ternary Bonsai 2 27B brings stronger reasoning, coding, vision, and agentic capability to the Bonsai series while preserving the deployment profile that defines it: a dramatically smaller memory footprint, high local throughput, and better energy efficiency.
基于 Qwen3.8 27B,Ternary Bonsai 2 27B 为 Bonsai 系列带来了更强的推理、编码、视觉和智能体能力,同时保留了定义该系列的部署特征:显著更小的内存占用、更高的本地吞吐量以及更好的能效。
Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight and a total model footprint of 5.9GB. The low-bit representation is applied end to end across the language model. It supports a 262K-token context window, multimodal text-and-image input, and is released under the Apache 2.0 license.
Ternary Bonsai 2 27B 使用三元 {-1, 0, +1} 权重配合 FP16 分组缩放,每个权重的有效位数为 1.76 位,模型总大小为 5.9GB。这种低位表示法贯穿整个语言模型端到端应用。它支持 262K token 的上下文窗口,支持多模态文本与图像输入,并以 Apache 2.0 许可证发布。
Against its full-precision counterpart, Ternary Bonsai 2 27B is more than 9x smaller while retaining 98.2% of aggregate benchmark performance. At this level of retention, compression becomes a deployment unlock: nearly the same capability, in a footprint that can run in far more places.
与全精度对应模型相比,Ternary Bonsai 2 27B 缩小了 9 倍以上,同时保留了 98.2% 的综合基准性能。在这个保留水平下,压缩成为部署的关键解锁因素:几乎相同的能力,却能在更多地方运行。
What changed from the first Bonsai 27B release
与首个 Bonsai 27B 版本相比的变化
Our first Bonsai 27B release was an important milestone, offering a practical way to run 27B-class intelligence on local devices. Bonsai 2 27B focuses on the next step: improving the model quality and runtime performance needed for real-world local applications. Compared with the previous Bonsai 27B generation, Bonsai 2 27B brings:
我们的首个 Bonsai 27B 版本是一个重要的里程碑,提供了一种在本地设备上运行 27B 级别智能的实用方法。Bonsai 2 27B 专注于下一步:提升现实世界本地应用所需模型质量和运行时性能。与上一代 Bonsai 27B 相比,Bonsai 2 27B 带来了:
- a stronger base model, Qwen3.8 27B
- higher aggregate capability retention of 98.2% against the full-precision model
- improved reasoning, coding, vision, and long-horizon agentic performance
- 更强大的基础模型 Qwen3.8 27B
- 相对于全精度模型,综合能力保留率提升至 98.2%
- 改进的推理、编码、视觉和长程智能体性能
Higher capability at the same deployment point
在相同的部署点实现更高的能力
Across a benchmark suite spanning reasoning, math, coding, instruction following, vision, and agentic tool use, Ternary Bonsai 2 27B scores 83.9, retaining 98.2% of Qwen3.8 27B’s aggregate performance.
在涵盖推理、数学、编码、指令遵循、视觉和智能体工具使用的基准测试套件中,Ternary Bonsai 2 27B 得分为 83.9,保留了 Qwen3.8 27B 综合性能的 98.2%。
| Capability | Ternary Bonsai 227B | Qwen3.827B | Qwen3.627B |
|---|---|---|---|
| Agentic & Tool Calling τ²-bench, BFCLv3 | 77.57 | 79.74 | 80.05 |
| Coding HumanEval+, LiveCodeBench v6, MBPP+, BigCodeBench | 81.58 | 82.17 | 82.57 |
| Instruction Following IFBench, IFEval | 82.66 | 81.25 | 74.53 |
| Knowledge & Reasoning MMLU-Redux, GPQA Diamond, AA-LCR | 83.95 | 86.66 | 84.71 |
| Math AIME 2026, AIME 2025, GSM8K, MATH-500 | 96.57 | 97.06 | 94.64 |
| Vision CharXiv, A-OKVQA, OmniDocBench v1.6, RealWorldQA, OCRBench v2 | 78.59 | 81.64 | 79.82 |
| Overall | 83.9 | 85.4 | 83.6 |
| 能力 | Ternary Bonsai 227B | Qwen3.827B | Qwen3.627B |
|---|---|---|---|
| 智能体与工具调用 τ²-bench, BFCLv3 | 77.57 | 79.74 | 80.05 |
| 编程 HumanEval+, LiveCodeBench v6, MBPP+, BigCodeBench | 81.58 | 82.17 | 82.57 |
| 指令遵循 IFBench, IFEval | 82.66 | 81.25 | 74.53 |
| 知识与推理 MMLU-Redux, GPQA Diamond, AA-LCR | 83.95 | 86.66 | 84.71 |
| 数学 AIME 2026, AIME 2025, GSM8K, MATH-500 | 96.57 | 97.06 | 94.64 |
| 视觉 CharXiv, A-OKVQA, OmniDocBench v1.6, RealWorldQA, OCRBench v2 | 78.59 | 81.64 | 79.82 |
| 综合 | 83.9 | 85.4 | 83.6 |
Figure I: Benchmark scores of Ternary Bonsai 2 27B (thinking mode) compared with the full-precision Qwen3.8 27B and Qwen3.6 27B baselines. Full per-benchmark results are in the whitepaper.
图 I:Ternary Bonsai 2 27B(思考模式)与全精度 Qwen3.8 27B 和 Qwen3.6 27B 基线的基准测试得分对比。各基准的详细结果见白皮书。
The key result is not only the aggregate score, but where the capability is retained. Coding agents, tool-use systems, multimodal workflows, and long-horizon tasks are particularly sensitive to model degradation because small errors can compound over many steps. Bonsai 2 27B preserves much of the full-precision model’s performance in exactly these areas while operating at a fraction of the memory footprint.
关键结果不仅在于总分,更在于能力的保留程度。编程智能体、工具使用系统、多模态工作流以及长周期任务对模型性能下降尤为敏感,因为微小的误差会在多个步骤中累积。Bonsai 2 27B 在恰好这些领域保留了全精度模型的大部分性能,同时内存占用仅为其一小部分。
Compared with the full-precision model and other low-bit alternatives, Bonsai 2 27B stands out as an outlier on intelligence density. Many low-bit alternatives become deployable only by giving up meaningful capability in coding, vision, or agentic tool use. Bonsai 2 27B pushes the frontier toward both higher capability and lower memory usage.
与全精度模型及其他低比特替代方案相比,Bonsai 2 27B 在智能密度方面表现突出。许多低比特替代方案只能通过牺牲编程、视觉或智能体工具使用方面的有意义能力来实现部署。Bonsai 2 27B 正推动前沿向更高能力和更低内存使用方向发展。
Figure II: Intelligence density (per GB) of Ternary Bonsai 2 27B compared to other models in the same parameter class.
图 II:Ternary Bonsai 2 27B 与其他同参数量级模型的智能密度(每 GB)。
Demo I: Coding agents with Cline, powered by Ternary Bonsai 2 27B on NVIDIA GeForce RTX 5090.
演示 I:由 NVIDIA GeForce RTX 5090 上的 Ternary Bonsai 2 27B 驱动的 Cline 编程智能体。
Demo II: Computer use powered by Ternary Bonsai 2 27B model on NVIDIA GeForce RTX 5090.
演示 II:由 NVIDIA GeForce RTX 5090 上的 Ternary Bonsai 2 27B 模型驱动的计算机操作。
‹›
‹›
With Bonsai 2 27B, local models can start to take on real knowledge work: coding-agent loops, computer-use workflows, private document analysis, multimodal debugging, and hybrid orchestration where local models handle sensitive or high-frequency tasks while escalating selectively to the cloud.
借助 Bonsai 2 27B,本地模型开始能够承担真正的知识工作:编程智能体循环、计算机操作工作流、私有文档分析、多模态调试,以及混合编排——其中本地模型处理敏感或高频任务,并选择性地将其他任务升级至云端。
Throughput and energy efficiency
吞吐量与能效
Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient than an 8B model running in full-precision.
Ternary Bonsai 2 27B 在 NVIDIA GeForce RTX 5090 上最高可达每秒 143 个 token,在 M5 Max 上为每秒 46.8 个 token。在 RTX 4090 上,Ternary Bonsai 2 27B 每 token 仅消耗 0.714 mWh,比以全精度运行的 8B 模型节能 40%。
For coding assistants, higher throughput means faster edit-debug loops. For multimodal agents, it means quicker iterations over screenshots, documents, and tool calls. For private local workflows, better energy efficiency means more useful inference on the same device, longer battery life, and a more realistic path to assistants that can stay available in the background without constantly calling the cloud.
对于编码助手而言,更高的吞吐量意味着更快的编辑-调试循环。对于多模态智能体,这意味着对截图、文档和工具调用的迭代速度更快。对于私有本地工作流,更好的能效意味着在同一设备上实现更有用的推理、更长的电池续航,以及让助手能够在后台持续可用而无需频繁调用云端的更现实路径。
Why this release matters
为何此次发布至关重要
Compared to Ternary Bonsai 27B, the new Ternary Bonsai 2 27B has closed the retention gap between the full precision model from 95% to over 98%. This is a significant improvement that makes the current release practically “lossless”. It further cements the notion that low-bit models can be the best way to deploy AI.
与 Ternary Bonsai 27B 相比,新的 Ternary Bonsai 2 27B 已将全精度模型从 95% 到超过 98% 之间的保留率差距缩小。这是一项显著的提升,使当前版本实际上达到了“无损”水平。这进一步巩固了低比特模型可能是部署 AI 的最佳方式这一观点。
That has implications well beyond local inference. Low-bit models can change the economics and architecture of AI systems across devices, workstations, and datacenters: fitting larger models into the same memory envelope, serving more users on the same hardware, reducing energy per inference, and enabling hybrid systems that dynamically decide what should run locally and what should run in the cloud.
其影响远远超出本地推理范畴。低比特模型可以改变跨设备、工作站和数据中心的 AI 系统的经济性和架构:在相同的内存包络中容纳更大的模型,在同一硬件上服务更多用户,降低每次推理的能耗,并支持混合系统动态决定哪些部分在本地运行,哪些部分在云端运行。
The question will increasingly be not just how capable a model is, but how much useful intelligence can be delivered within a given memory, compute, and power budget. If capability can continue to scale while those requirements fall dramatically, the deployment envelope for future models expands across the stack: from personal devices to large-scale datacenters.
未来的问题将不再仅仅是模型具备多少能力,而是在给定的内存、计算和功耗预算内能交付多少有用的智能。如果能力能够继续扩展,而这些要求却大幅下降,那么未来模型的部署范围将在整个技术栈中扩大:从个人设备到大规模数据中心。
Platform Coverage
平台覆盖
Bonsai 2 27B runs on NVIDIA GPUs via CUDA and on Apple devices (Mac, iPhone, iPad) via MLX, through custom low-bit kernels. Model weights are available today under the Apache 2.0 License.
Bonsai 2 27B 通过自定义低比特内核,在 NVIDIA GPU 上通过 CUDA 运行,并在 Apple 设备(Mac、iPhone、iPad)上通过 MLX 运行。模型权重目前已根据 Apache 2.0 许可证提供。
Full technical details of our compression, evaluation, and benchmarking processes are available in our whitepaper.
我们压缩、评估和基准测试流程的完整技术细节可在我们的白皮书中找到。
Work with Us
与我们合作
We work with teams to tailor Bonsai models to their applications, from post-training on domain-specific data to optimizing inference for target hardware. If you’re building AI products with tight memory, latency, or power requirements, we’d love to explore how Bonsai can help. Reach out at [email protected].
我们与团队合作,将 Bonsai 模型定制应用于他们的场景,从针对特定领域数据的后训练到为目标硬件优化推理。如果您正在构建对内存、延迟或功耗有严格要求的 AI 产品,我们很乐意探讨 Bonsai 如何提供帮助。请联系 [email protected]。
Join Us
加入我们
PrismML emerged from a team of Caltech researchers and was founded with support from Khosla Ventures, Cerberus, and Google, with continuing support from Samsung. We've spent years tackling one of the field's hardest problems: compressing neural networks without sacrificing their reasoning ability.
PrismML 由加州理工学院的研究团队创立,并获得 Khosla Ventures、Cerberus 和 Google 的支持,同时得到三星的持续支持。多年来,我们致力于解决该领域最棘手的问题之一:在不牺牲推理能力的情况下压缩神经网络。
If you want to help build the next generation of state-of-the-art AI, we'd love to hear from you. Check out our careers page.
如果您希望参与构建下一代最先进的 AI,我们非常期待与您联系。请查看我们的招聘页面。
Back to all posts
返回所有文章
Announcing Bonsai 27B: The First 27B-Class Model to Run on a Phone
宣布 Bonsai 27B:首款可在手机上运行的 27B 级模型
July 14, 2026
2026年7月14日
Today we're announcing Bonsai 27B, our multimodal flagship: ternary at 5.9GB for laptops, 1-bit at 3.9GB for an iPhone 17 Pro, with a 262K-token context.
今天我们宣布推出 Bonsai 27B,这是我们的多模态旗舰产品:笔记本电脑端采用三元模型,大小为 5.9GB;iPhone 17 Pro 端采用 1-bit 模型,大小为 3.9GB,支持 262K token 上下文。
PrismML Launches Bonsai 2 27B, Its Most Capable Model Yet
PrismML 发布 Bonsai 2 27B,其迄今最强大的模型
September 17, 2026
2026年9月17日
New flagship model brings 27B-class reasoning, coding, vision, and agentic capability into a dramatically smaller, faster deployment footprint
全新旗舰模型将 27B 级的推理、编码、视觉和智能体能力带入大幅更小、更快的部署规模
Thanks, we’ll keep you posted!
谢谢,我们会及时向您通报!
Something went wrong.
发生错误。
Resources
资源
Demo
演示
Whitepaper
白皮书
Docs
文档
Models
模型
Hugging Face
Hugging Face
GitHub
GitHub
Follow
关注
X
X
Discord
Discord
Contact us
联系我们
Have a question, partnership idea, or a project that needs efficient intelligence? Reach out—we’d love to hear from you.
有问题、合作想法或需要高效智能的项目?请联系我们要——我们非常期待听到您的声音。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力