小模型时代已至:性能逼近大模型,成本大幅降低
Small Models Have Arrived
小型模型已经到来
AUG 26, 2026
2026年8月26日
For the past few weeks, I've been playing with gpt-5.6-luna. It is shockingly capable, fast, and smart. I regularly see it do ~100 tps, and rip around my codebase, email, and knowledge base.
过去几周,我一直在玩gpt-5.6-luna。它能力惊人,速度快,又聪明。我经常看到它以约100 tps的速度运行,并快速浏览我的代码库、电子邮件和知识库。
Of course, the biggest thing with luna is the cost. I've tried running some fairly complicated research threads, and it's pretty tough to run up a large bill. Even having it search across thousands of emails, I end up with an API cost in the tens of cents.
当然,luna最大的特点是成本。我尝试运行一些相当复杂的研究线程,但很难产生高额账单。即使让它搜索数千封电子邮件,API成本也只有几美分。
Courtesy of artificialanalysis.ai
数据来源:artificialanalysis.ai
With GLM 5.3, we even have a new option at the Pareto frontier.
有了GLM 5.3,我们在帕累托前沿甚至有了新的选择。
When doing coding work, I almost always reach for the most expensive and capable models (Fable 5, 5.6 Sol). So it's been easy to miss the progress the small fast models have made.
在做编码工作时,我几乎总是选择最昂贵、功能最强的模型(如Fable 5、5.6 Sol)。因此,很容易忽视小型快速模型所取得的进步。
One thing a few investors I've talked with have mentioned: "It's weird we're not seeing more consumer AI companies. Why is that?"
我交谈过的几位投资者提到的一件事:“奇怪的是,我们没有看到更多的消费级AI公司。这是为什么?”
There's a straightforward answer: token costs.
答案很简单:令牌成本。
In the times before AI, the playbook for big consumer apps looked like this...
在AI出现之前,大型消费应用的标准做法是这样的...
- create some sort of compelling website which is fairly cheap to run
- attract a bunch of users (typically with some virality)
- raise money, scale to more users
- create an ads marketplace
- 创建某种有吸引力的网站,运行成本相当低
- 吸引大量用户(通常通过某种病毒式传播)
- 筹集资金,扩展到更多用户
- 创建广告市场
This roughly describes most of the big consumer companies (Google, Facebook, Snapchat, etc.).1
这大致描述了许多大型消费公司(如谷歌、Facebook、Snapchat等)的情况。
But what if you want to add AI to your product? Well, now you have some real inference costs on every request! Suddenly the amount of capital required increases dramatically.
但如果你想在产品中加入AI呢?那么,现在每次请求都会产生实际的推理成本!突然之间,所需资本量急剧增加。
A pet eval of mine is to build a daily news site, personalized to me:
我个人的一个评估是建立一个每日新闻网站,为我个性化定制:
research @calvinfo on the internet. figure out what news they might like. build a micro-site with today's top stories, personalized for them. search hn, reddit, twitter, etc.
在互联网上研究@calvinfo。弄清楚他们可能喜欢什么新闻。构建一个微型网站,展示今天的头条新闻,为他们个性化定制。搜索hn、reddit、twitter等。
With the previous generation of models (Sonnet class), you'd spend ~$1 to get anywhere. Charging $30/mo is untenable for a consumer app. There's obviously a lot we can optimize here, but if you're charging what the WSJ or The Economist charges, you'd better be delivering similar value.
使用上一代模型(如Sonnet级别),你大约要花费1美元才能有所进展。对于消费应用来说,每月收费30美元是不可行的。显然,这里有很多可以优化的地方,但如果你收费与WSJ或《经济学人》相当,你最好提供类似的价值。
But looking at luna, the results are pretty decent, and the average cost is ~$0.10. Now we're talking!
但看看luna,结果相当不错,平均成本约为0.10美元。这才像话!
Where I think this gets even more interesting is in the world of business.
我认为这更有趣的地方在于商业领域。
My Segment co-founder Peter and I were recently comparing notes on a hike. Across his various startups, Peter has seen two kinds of work:
我和我的Segment联合创始人Peter最近在一次徒步中交流了笔记。在他的各种初创公司中,Peter看到了两种工作:
- the "IQ 180" work. some mad scientist genius type comes up with some crazy solution you've never thought of.
- the "token spewer" work. being ultra responsive, pushing the ball forward across dozens of different fronts.
- “IQ 180”型工作。某个疯狂科学家天才类型的人提出了你从未想过的疯狂解决方案。
- “代币喷射器”式的工作。超快响应,在数十条不同战线上推进事务。
Peter runs multiple companies. Beyond Segment, he's raised $100m+ for Charm Industrial, and just recently closed a Series A for Revoy. He's incredibly organized and efficient with his time.
彼得经营着多家公司。除了Segment,他还为Charm Industrial筹集了超过1亿美元,最近刚为Revoy完成了A轮融资。他组织能力极强,时间利用效率极高。
And yet, Peter mentioned that ~95% of the work he does falls into bucket 2. It's hopping on calls. Nudging people. Blocking and tackling.
然而,彼得提到他约95%的工作属于第二类。就是打电话、催促他人、处理琐事和防守反击。
To be clear, Peter says his companies would be dead-in-the-water today without an IQ 180 technical mind solving the deep problems. Just that most of his work falls in bucket 2.2
需要说明的是,彼得表示如果没有IQ 180的技术头脑解决深层问题,他的公司今天就会陷入困境。只是他的大部分工作属于第二类。
I think demand for "frontier-level" models is going to keep compounding. Especially for fields that require novel breakthroughs or discovery (engineering, hard science, model training).
我认为对“前沿级”模型的需求将持续增长。尤其是在需要新颖突破或发现的领域(工程、硬科学、模型训练)。
But I also think the demand for "fast/cheap/good-enough" models is just about to take off.
但我也认为对“快速/廉价/够用”模型的需求即将爆发。
Think of the people you interact with on a daily basis: coworkers, vendors, and customers. Nine times out of ten, you want someone who is super responsive, and just handles things for you. Most of the "human tokens" at companies today are spent this way — hiring skews heavily toward the fast/cheap/good-enough archetype.
想想你日常接触的人:同事、供应商和客户。十有八九,你希望对方反应迅速,直接帮你把事情搞定。如今公司里大部分“人力代币”都花在这方面——招聘严重偏向快速/廉价/够用型人才。
There's a lot of work that needs to happen to make fast/cheap/good-enough models a reality for business. New harnesses, prompt injection safety, roles, and permissions. But I'm confident we'll figure that out.
要让快速/廉价/够用模型在商业中成为现实,还有很多工作要做。新的框架、提示注入安全、角色和权限。但我相信我们一定能解决这些问题。
If you're also experimenting with making small models useful, please drop me a line.
如果你也在尝试让小型模型变得有用,请随时联系我。
Footnotes
脚注
- Amazon and Netflix are the notable exceptions ↩
- Peter is also being modest here. He's sharp as a tack. ↩
- 亚马逊和Netflix是显著的例外 ↩
- 彼得在这里也很谦虚。他思维敏捷,非常聪明。 ↩
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力