跳到主内容
精选75Tomasz Tunguz(RSS)战略与拆解

谁在买SOTA模型?数据揭示性价比才是主流

Honestly, Who Buys SOTA?

原文
推荐理由

给做AI产品与采购的读者:用数据说明多数场景下性价比模型已够用,并给出价格弹性与份额转移的具体数字,可直接用于模型选型与成本决策。

State of the art models are two-thirds smarter than they were last November. The frenetic pace of improvement is sustained, two new models every three days. 1

最先进的模型比去年十一月聪明了三分之二。改进的速度疯狂持续,每三天就有两个新模型问世。1

But 84% of tokens on OpenRouter aren’t state of the art. 2 3

但 OpenRouter 上 84% 的 token 并非来自最先进的模型。2 3

In fact, the six models users choose to generate the supermajority of those tokens deliver about 77% of the performance of the frontier. They cost 2.5% of what Claude Fable 5 does. 2

事实上,用户选择生成绝大多数 token 的六个模型,其性能约为前沿模型的 77%。它们的成本仅为 Claude Fable 5 的 2.5%。2

The index keeps jumping. Large gains of three to five Artificial Analysis points land about every quarter. Smaller steps fill the gaps.

指数持续跃升。大约每季度就会出现三到五个 Artificial Analysis 积分的大幅增长。较小的步骤填补了空白。

Six models carry 80% of volume in the week of August 10. Their blended price is $0.50 per million tokens against Fable 5 at $20.

在 8 月 10 日那一周,六个模型承载了 80% 的流量。它们的混合价格为每百万 token 0.50 美元,而 Fable 5 为 20 美元。

Ramp’s data shows buyers are price-elastic. Fable 5 at about $10/m tokens captured 6% of Anthropic tokens & 11% of Anthropic spend a month after launch. GPT-5.6 Sol, OpenAI’s priciest mainline tier, held about a quarter of OpenAI tokens. 4

Ramp 的数据显示买家对价格敏感。Fable 5 定价约为每百万 token 10 美元,在发布一个月后占据了 Anthropic token 的 6% 和 Anthropic 支出的 11%。GPT-5.6 Sol 是 OpenAI 最昂贵的主流层级,占据了 OpenAI token 的约四分之一。4

Fable 5 generated roughly 75% as much model-attributed revenue as GPT-5.6 Sol in July, despite being substantially more expensive.

7 月份,Fable 5 产生的模型归属收入约为 GPT-5.6 Sol 的 75%,尽管其价格要高得多。

Each new state of the art release should move less share than the one before it.

每一次新的最先进模型发布,其市场份额的变动应小于前一次。

Enterprises will consolidate spend. Contracts concentrate on one or two vendors, just like in the cloud era, & once a model clears a high-value job the workload stays.

企业将整合支出。合同集中在少数几个供应商上,就像云时代一样,一旦某个模型通过了高价值任务的考验,工作负载就会保留下来。

Performance is already good enough at a meaningful discount. The gap keeps closing from below. The best open-weight model reached 80% of the frontier score by May, up from 48% a year earlier. 2

性能已经足够好,而且价格有显著折扣。差距从下方不断缩小。最佳开源权重模型在 5 月份达到了前沿模型分数的 80%,而一年前仅为 48%。2

Application deployment is the other story. More of our portfolio companies & startups default to smaller models, fine-tuned models, & open source. They are optimizing against a different Pareto frontier, price over performance.

应用部署是另一回事。我们更多的投资组合公司和初创公司默认使用较小的模型、微调模型和开源模型。他们正在针对不同的帕累托前沿进行优化,即价格优先于性能。

If share stops shifting & good enough stays good enough, the economics of SOTA change. A nine-figure training run has to win share to pay for itself, & that bar will rise with time.

如果市场份额停止转移,且“足够好”仍然足够好,那么最先进模型的经济性就会改变。一次九位数的训练运行必须赢得市场份额才能收回成本,而这个门槛会随着时间推移而提高。

The title is flippant. Plenty buy SOTA, & for good reason. Software engineering architecture & security design are the clearest cases, where the best available model earns its price.

这个标题有点轻率。很多人购买最先进的模型,而且理由充分。软件工程架构和安全设计是最明显的例子,在这些领域,最好的可用模型物有所值。

But the open data we do have suggests the frontier that matters is the other one.

但我们确实拥有的公开数据表明,真正重要的前沿是另一个。

  • Artificial Analysis model catalog & Intelligence Index. Major-lab monthly release counts & frontier path. Sample starts 2025-11-01. Release-rate trend flat. Median gap between large (≥3 pt) frontier steps about 3.5 months. Intelligence Index ↩︎
  • State of the art means the single best Artificial Analysis score available in a given week; a model counts as near it when the score sits within 10% of that week’s best named model. OpenRouter weekly named top models joined to Artificial Analysis scores, the head of the OpenRouter carousel rather than every API. Share series, weeks 2025-11-03 through 2026-05-25 (n=30), named only, Others excluded. First vs last thirteen weeks about 17.5% vs 14.6% near the frontier (~82-85% outside). Concentration snapshot, week of 2026-08-10, models covering the first ~80% of named tokens. Token-weighted Artificial Analysis about 23% behind global catalog state of the art (~77% of frontier quality) & about 10% behind the best model on that OpenRouter list. Blended basket about $0.50/m tokens vs Fable 5 at $20/m (~40x). May 2026 historical check, about 16% behind the local list leader. Best open-weight model 47.5% of frontier score in the first thirteen weeks vs 70.9% in the last thirteen; single best week 2026-05-25, DeepSeek V4 Pro at 45.27 vs frontier 56.31 (80.4%). OpenRouter rankings ↩︎ ↩︎ ↩︎
  • These data sources don’t capture the first-party clouds, OpenAI, Anthropic & Google’s own services. Frontier traffic running on native APIs never enters the OpenRouter rankings, so there’s a bias to the data. ↩︎
  • Ramp Economics Lab, AI Index August 2026 (Fable 5 uptake). econlab.substack.com/p/ai-index-august-2026 ↩︎
  • Artificial Analysis 模型目录和智能指数。主要实验室每月发布数量和前沿路径。样本开始于 2025 年 11 月 1 日。发布率趋势平稳。前沿大步(≥3 分)之间的中位间隔约为 3.5 个月。智能指数 ↩︎
  • 最先进水平指的是在给定一周内可获得的最佳单一人工智能分析分数;当分数在该周最佳命名模型的10%以内时,该模型被视为接近最先进水平。OpenRouter每周命名的顶级模型与人工智能分析分数相结合,取OpenRouter轮播图的头部而非每个API。分享系列,周次从2025-11-03至2026-05-25(n=30),仅命名模型,排除其他。前十三周与后十三周相比,接近前沿的比例约为17.5%对14.6%(约82-85%在外)。集中度快照,2026-08-10那一周,模型覆盖了命名代币的前约80%。按代币加权的人工智能分析分数比全球目录最先进水平落后约23%(约为前沿质量的77%),比OpenRouter列表上的最佳模型落后约10%。混合篮子约为每百万代币0.50美元,而Fable 5为每百万代币20美元(约40倍)。2026年5月的历史检查,比本地列表领先者落后约16%。最佳开放权重模型在前十三周的前沿分数为47.5%,而在后十三周为70.9%;单周最佳为2026-05-25,DeepSeek V4 Pro得分为45.27,前沿为56.31(80.4%)。OpenRouter排名 ↩︎ ↩︎ ↩︎
  • 这些数据源不涵盖第一方云服务,如OpenAI、Anthropic和Google自己的服务。运行在原生API上的前沿流量从未进入OpenRouter排名,因此数据存在偏差。 ↩︎
  • Ramp经济学实验室,2026年8月AI指数(Fable 5采用情况)。econlab.substack.com/p/ai-index-august-2026 ↩︎

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近