本地AI模型能效提升:路由策略与成本优化
Mainframes became personal. So will your data center.
提供了具体的架构选型建议(本地+路由)和量化的成本/能效收益数据,适合正在评估AI部署方案的团队参考。
Local AI models can already answer 89% of everyday chat & reasoning queries as well as a frontier cloud model, a result that holds broadly across more than a million real queries & 20+ local models tested.1
本地 AI 模型已经能够以与前沿云模型相当的水平回答 89% 的日常聊天和推理查询,这一结果在超过一百万个真实查询以及 20 多种本地模型的测试中普遍成立。1
That means we’re generating more intelligence per watt of electricity : more work from the same number of electrons.1
这意味着我们每瓦特电力产生的智能更多:相同数量的电子完成了更多工作。1
Tracking computing efficiency over time is not new. Koomey’s law found that computing power per watt doubled roughly every 1.5 years for decades, a trend that shrank the power of a mainframe into a laptop’s chassis.23
长期追踪计算效率并非新鲜事。Koomey 定律发现,数十年来每瓦特的计算能力大约每 1.5 年翻一番,这一趋势将大型机的功耗缩小到了笔记本电脑的机箱内。23
Just as performance-per-watt guided the mainframe-to-PC transition, intelligence-per-watt will guide AI’s transition to the edge.
正如每瓦性能指导了从大型机到个人电脑的过渡,每瓦智能也将指导 AI 向边缘计算的过渡。
— Jon Saad-Falcon, Avanika Narayan, et al., “Intelligence per Watt”
——Jon Saad-Falcon、Avanika Narayan 等人,《每瓦智能》
GPUs follow a more languid curve today, doubling efficiency roughly every 2.7 years over the last 15 years, not every 1.5 years.4
如今 GPU 遵循一条更为平缓的曲线,在过去 15 年中效率大约每 2.7 年翻一番,而非每 1.5 年。4
The impact is no less impressive : the best local model’s win/tie rate against a frontier model, rose from 23.2% in 2023 to 71.3% in 2025, adding roughly 20 percentage points a year. In 2026, a local model selected for the task by another local AI bumps that number to nearly 90%.5
其影响同样令人印象深刻:最佳本地模型对阵前沿模型的胜率/平局率从 2023 年的 23.2% 上升至 2025 年的 71.3%,每年增加约 20 个百分点。在 2026 年,由另一个本地 AI 为该任务选择的本地模型将该数字推高至接近 90%。5
Efficiency improved alongside the trend line : intelligence-per-watt rose 5.3x over the same period, split into a 3.1x gain from better models & a 1.7x gain from better chips. Computers are faster, models are smarter. And the end user benefits.
效率随趋势线同步提升:同一时期内每瓦智能提升了 5.3 倍,其中模型改进贡献了 3.1 倍的增益,芯片改进贡献了 1.7 倍的增益。计算机更快了,模型更聪明了,最终用户从中受益。
5
5
Cloud remains essential for long multi-step reasoning, the hardest technical domains, & workloads where scale & parallelization matter. Cloud hardware still holds an edge over local hardware in many use cases. Cloud inference delivers a 40% energy efficiency gain relative to local models.6 Datacenters batch queries, a trick local hardware serving one user at a time cannot use, yet.
云端对于长多步推理、最困难的技术领域以及规模和平行化至关重要的负载仍然不可或缺。在许多用例中,云端硬件仍优于本地硬件。相对于本地模型,云端推理带来了 40% 的能效增益。6 数据中心对查询进行批处理,这是一种目前仅一次服务一个用户的本地硬件尚无法使用的技巧。
But for much of everyday knowledge work, there is no reason to send the query to the data center at all. The combination of local models plus a router is sufficient for the supermajority of work.7
但对于大量日常知识型工作而言,完全没有理由将查询发送到数据中心。本地模型加上路由器的组合足以应对绝大多数工作。7
It also cuts energy 80%, compute 77%, & cost 74% against an all-cloud baseline.
与全云基线相比,它还能减少 80% 的能耗、77% 的计算量和 74% 的成本。
Mainframes became personal. So will your data center.
大型机变成了个人电脑。你的数据中心也将如此。
- Jon Saad-Falcon, Avanika Narayan, et al., “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI,” Stanford University & Together AI, November 2025. arXiv:2511.07885. See also the Stanford Hazy Research overview. ↩︎ ↩︎
- Koomey’s law, Wikipedia. ↩︎
- Watts measure power, the rate energy is drawn, & joules measure the energy itself; Koomey’s original metric was computations per joule, but the underlying trend is the same one intelligence per watt now tracks for AI models. ↩︎
- Anson Ho, Ege Erdil & Tamay Besiroglu, “Limits to the Energy Efficiency of CMOS Microprocessors,” 2023. arXiv:2312.08595. ↩︎
- 71.3% & 89% are two different measurements from the same 2025 data, not the same number at different times. 71.3% is the best single local model’s win/tie rate against a frontier model, the trend line this chart plots (23.2% in 2023, 48.7% in 2024, 71.3% in 2025). 89% is a separate, higher ceiling: routing each query to whichever of the 20+ local models tested handles it best beats any single model by 16.3 to 28.8 percentage points. That gain is a selection effect: the paper notes local routing draws from 20+ diverse models versus three frontier cloud models, so on some benchmarks the best-of-local ensemble even surpasses best-of-cloud. More candidates to choose from, not just smarter routing, is what raises accuracy. Both figures come from the same study & neither supersedes the other. ↩︎ ↩︎
- Jon Saad-Falcon, Avanika Narayan, et al., “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI,” Stanford University & Together AI, November 2025. Cloud accelerators deliver at least 1.4x higher intelligence-per-watt than local chips running the same models, roughly a 40% efficiency premium. arXiv:2511.07885. ↩︎
- Most AI Work Can Wait, tomtunguz.com. ↩︎
- Jon Saad-Falcon、Avanika Narayan 等人,《每瓦智能:衡量本地 AI 的智能效率》,斯坦福大学与 Together AI,2025 年 11 月。arXiv:2511.07885。另见斯坦福 Hazy Research 概述。↩︎ ↩︎
- Koomey 定律,维基百科。↩︎
- 瓦特衡量功率,即能量消耗的速率;焦耳衡量能量本身。Koomey 最初的指标是每焦耳的计算量,但底层趋势相同,现在用“每瓦智力”来追踪 AI 模型。↩︎
- Anson Ho、Ege Erdil 与 Tamay Besiroglu,《CMOS 微处理器能效极限》,2023 年。arXiv:2312.08595。↩︎
- 71.3% 和 89% 是来自同一份 2025 年数据的两种不同测量结果,并非同一数字在不同时间的数值。71.3% 是单个最佳本地模型在与前沿模型对抗时的胜率/平局率,也是该图表绘制的趋势线(2023 年为 23.2%,2024 年为 48.7%,2025 年为 71.3%)。89% 是一个独立的更高上限:将每个查询路由到测试过的 20 多个本地模型中处理效果最好的一个,比任何单一模型的准确率高出 16.3 至 28.8 个百分点。这种增益源于选择效应:论文指出,本地路由从 20 多个多样化模型中进行选择,而云端前沿模型仅有三个,因此在某些基准测试中,最佳本地集成甚至优于最佳云端集成。可选择的候选者更多,而不仅仅是更智能的路由机制,才提升了准确率。这两个数据均出自同一项研究,且互不替代。↩︎ ↩︎
- Jon Saad-Falcon、Avanika Narayan 等人,《每瓦智力:衡量本地 AI 的智力效率》,斯坦福大学与 Together AI,2025 年 11 月。云端加速器在运行相同模型时,其每瓦智力至少比本地芯片高 1.4 倍,效率溢价约为 40%。arXiv:2511.07885。↩︎
- 大多数 AI 工作可以等待,tomtunguz.com。↩︎
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力