跳到主内容
@wquguru
精选88Avi Chawla模型发布/更新

开源推理模型实际成本远超报价,存在“过度思考税”

This chart will worry open-source labs!

原文
发到 X
推荐理由

揭示了开源推理模型在计费效率上的结构性劣势,直接影响选型与部署成本核算,建议关注其提供的交互式仪表盘进行量化对比。

This chart will worry open-source labs!

这张图表会让开源实验室感到担忧!

Researchers just measured the actual dollars models consume per solved task across dozens of LLMs and 16 benchmarks, then compared that against quoted per-token pricing.

研究人员刚刚测量了数十种大语言模型在解决任务时实际消耗的美元金额,并基于 16 个基准测试进行了对比,随后将其与按 token 报价的定价进行了比较。

Sorted by how far real cost exceeds quoted pricing, all 8 models at the worst end are open-weight, coming from Kimi, DeepSeek, Qwen, and GLM.

按实际成本超出报价的程度排序,最糟糕一端的 8 个模型均为开源权重模型,分别来自 Kimi、DeepSeek、Qwen 和 GLM。

The first closed model appears tenth, and on the favorable end, the first open-weight model shows up ninth.

第一个闭源模型排在第十位,而在有利的一端,第一个开源权重模型出现在第九位。

A pricing page quotes dollars per token, but it cannot quote how many tokens a model needs, and that second number is where open-weight reasoning models spend heavily.

定价页面按 token 报价美元,但无法说明模型需要多少 token,而第二个数字正是开源权重推理模型大量消耗的地方。

R1-style models are trained with rewards on the final answer only.

R1 风格的模型仅在最终答案上通过奖励进行训练。

A longer chain of thought raises the odds of hitting a correct step, so training reinforces long chains, and every one of those thinking tokens is billed as output.

更长的思维链提高了命中正确步骤的概率,因此训练会强化长思维链,而这些思考过程中的每一个 token 都会作为输出被计费。

Closed labs appear to tune this harder, with length-aware rewards and effort controls that spend tokens only where the task needs them.

闭源实验室似乎对此进行了更深入的优化,采用长度感知的奖励和努力控制机制,仅在任务需要的地方消耗 token。

A recent benchmark, OckBench, measured the same efficiency gap between open-weight and proprietary models and called it the overthinking tax.

最近的基准测试 OckBench 测量了开源权重模型与专有模型之间的同一效率差距,并将其称为“过度思考税”。

Tokenizers widen the gap further, since every lab counts tokens differently. Tibo from OpenAI made that half of the argument recently, that a lower price per token doesn't guarantee a lower bill.

分词器进一步拉大了这一差距,因为每家实验室对 token 的计算方式各不相同。OpenAI 的 Tibo 最近指出这一点,即较低的单价并不保证账单更低。

Open-weight models still make sense for self-hosting, fine-tuning, and data control. But pricing is not the right way to compare them.

对于自托管、微调和数据控制而言,开源权重模型仍然有意义。但定价并不是比较它们的合适方式。

All of this is now documented in AI Frontier. It's an interactive dashboard by Martian, and the method behind it is a peer-reviewed ICLR 2026 paper.

所有这些内容现已记录在 AI Frontier 中。这是 Martian 制作的一个交互式仪表板,其背后的方法是一篇经过同行评审的 ICLR 2026 论文。

The screenshot below depicts the published dashboard, and I worked with the Martian team today to share it.

下面的截图展示了已发布的仪表板,我今天与 Martian 团队一起分享了它。

More details in the quote below.

更多细节见下方的引用内容。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近