递归自我改进的极限与开放权重模型的商业崛起
🔮 Unbounded self-improvement and its limits #599
Good morning!
早上好!
The limits to recursion
递归的极限
What are the conditions under which AI systems could undergo recursive self-improvement (RSI) — and how long might that last? Cards on the table, I’m not wildly excited by the theory of unending accelerating recursive self-improvement for the simple theoretical issue of control and alignment. But I also think it’s not likely for practical and theoretical reasons.
在什么条件下,AI系统可能经历递归自我改进(RSI)——这又能持续多久?坦白说,对于无限加速的递归自我改进理论,我并不特别热衷,原因在于控制和校准方面的简单理论问题。但我也认为,出于实践和理论上的原因,这不太可能发生。
Now philosopher, Toby Ord, has done the heavy lifting for me. He argues that a key gating factor to RSI is the generation time: how long it takes of an AI system to go through a single loop of improvement (where it helps design and train its successor).
现在,哲学家托比·奥德已经为我做了繁重的工作。他认为,RSI的一个关键限制因素是生成时间:即AI系统完成一轮改进(在此过程中它帮助设计和训练其继任者)所需的时间。
Ord concludes that while it is mathematically possible for extreme RSI, intelligence rising without bound, the conditions are unbearably difficult to achieve. The key question is whether the entire research to training to development cycle can shrink to zero or not. Ord reckons unlikely, I too don’t believe it is possible.
奥德得出结论,虽然极端RSI在数学上是可能的,即智能无界上升,但实现条件极其困难。关键问题是,从研究到训练再到开发的整个周期能否缩短至零。奥德认为不太可能,我也不相信这是可能的。
Generation time can’t get to zero because real-life intrudes: experiments take time; training runs take time; making new chips take time… lots of things take time. It might still feel fast, but it wouldn’t race to infinity.
生成时间无法降至零,因为现实生活介入:实验需要时间;训练运行需要时间;制造新芯片需要时间……很多事情都需要时间。它可能仍然感觉很快,但不会飞速冲向无限。
Eventually, physics intrudes too: the speed of light limits communication speed; the Bekenstein bound limits the information contained within a finite bit of space; and Landauer imposes an energy tax on irreversible computation.1
最终,物理学也会介入:光速限制了通信速度;贝肯斯坦界限制了有限空间内包含的信息量;而兰道尔原理对不可逆计算施加了能量代价。
The Universe, it seems, agrees with me. Unbounded RSI has its limits.
宇宙似乎也同意我的观点。无界的RSI有其极限。
You can read Ord’s paper here. Premium members can explore a plain English interactive version too.
你可以在这里阅读奥德的论文。高级会员还可以探索一个简明英语的互动版本。
Subscribe now
立即订阅
The market is the instrument
市场是工具
Open-weight models are growing in popularity in the business world: their token share at Vercel hit a single-day record of 62%, up from 28% two months earlier. Some Western firms are even moving workloads to Chinese open weights — Thomson Reuters has developed its first in-house model based on Qwen to cut costs. You can tune cost-effective open weights to match, or sometimes beat, frontier performance on the tasks that matter to you. Take Bridgewater: working with Thinking Machines, it fine-tuned an open Qwen model on expert-labeled data and beat every frontier model it tested on its internal information-filtering tasks: roughly 30% fewer errors than the best closed model, at one-fourteenth of the inference cost. Trainloop, which I am an investor in, does something similar, using the tiny Qwen 3.7-27b model, and can outperform GPT 5.6 Sol on specific fine-tuned tasks at a fraction of the cost.
开源权重模型在商业界越来越受欢迎:其在Vercel的令牌份额达到了62%的单日纪录,较两个月前的28%有所上升。一些西方公司甚至将工作负载迁移至中国的开源权重模型——汤森路透已基于Qwen开发了其首个内部模型以降低成本。您可以调整成本效益高的开源权重模型,以匹配甚至在某些对您重要的任务上超越前沿性能。以桥水基金为例:它与Thinking Machines合作,在专家标注的数据上微调了一个开源的Qwen模型,在其内部信息过滤任务上击败了所有测试过的前沿模型:比最佳封闭模型错误率低约30%,而推理成本仅为后者的十四分之一。我作为投资者的Trainloop也做了类似的事情,使用小巧的Qwen 3.7-27b模型,在特定微调任务上能以极低的成本超越GPT 5.6 Sol。
The results that Trainloop is getting are pretty impressive, seeing as they are based on a pocket model—a 27b model will even fit on a desktop Mac.2
Trainloop取得的结果相当令人印象深刻,考虑到它们基于的是一个袖珍模型——一个27b的模型甚至能装进台式Mac。2
Openweight models are, of course, getting better and better. Z.ai GLM 5.3, released this week, completely reshapes the cost-performance Pareto frontier. Of course, it isn’t a small model, but smaller distillations will emerge from it.
当然,开源权重模型正在变得越来越好。本周发布的Z.ai GLM 5.3完全重塑了成本性能的帕累托前沿。当然,它不是一个小模型,但更小的蒸馏版本将会从中产生。
All of this speaks to a welcome competition in AI provision. Clearly, firms could move focused workloads onto the most-performant, fine-tuned small models they can. Where possible, they might choose large, generally capable open models. But the appeal of being at the frontier, which is more than just model performance—it is service guarantees, harness quality, reliability, and a host of other requirements—still drives significant business for Anthropic and OpenAI.
这一切都体现了AI供应领域令人欢迎的竞争。显然,公司可以将重点的工作负载转移到他们能找到的性能最优、微调得当的小模型上。在可能的情况下,他们可能会选择大型、通用能力强的开源模型。但处于前沿的吸引力——这不仅仅是模型性能,还包括服务保证、工具质量、可靠性以及一系列其他要求——仍然为Anthropic和OpenAI带来了大量业务。
We don’t think it has much impact on the question of whether revenues flowing into the industry will materially change. For one thing, we don’t have a counterfactual to test against. But more importantly, every open model still involves paying inference providers. We’ll be looking at this question in more detail in the comings weeks.
我们认为这对流入行业的收入是否会实质性变化影响不大。一方面,我们没有反事实可以测试。但更重要的是,每个开源模型仍然涉及向推理提供商付费。我们将在未来几周更详细地研究这个问题。
A spicy entry to the compute ecosystem
计算生态系统中的一个辛辣新成员
Recursive self-improvement may be entering the compute realm. OpenAI’s new chip, ‘Jalapeño’, was designed with a heavy helping hand from the company’s own models, which helped write kernels and cut roughly 10% from one of the chip’s main compute blocks. In around 16 months from first hire to tape-out, OpenAI has built a chip that beats comparable Nvidia silicon by 1.5–1.9x on tokens per megawatt at peak throughput. This suggests frontier models can compress the design cycle for competitive silicon.
递归自我改进可能正在进入计算领域。OpenAI 的新芯片“Jalapeño”在设计时得到了公司自身模型的大力协助,这些模型帮助编写了内核,并从芯片的一个主要计算模块中削减了约10%的规模。从首次招聘到流片,OpenAI 在约16个月内构建了一款芯片,在峰值吞吐量下,每兆瓦的令牌数比同类英伟达硅片高出1.5至1.9倍。这表明前沿模型可以压缩竞争性硅片的设计周期。
AI will result in far greater heterogeneity in chip architectures than we saw in prior computing markets. Personal computers battled between the x86 standard and the Motorola 68x, before today’s duopoly of Intel and Apple silicon. Different uses for AI will need compute optimised for intelligence, latency, power consumption, training, and inference. This creates lots of room for specialist firms. One example is that ChatGPT’s fast response mode is powered by Cerebras’ low-latency silicon. Another is Fractile, where I am an investor, which has a deal with Anthropic for its low-latency inferencing chips.
人工智能将导致芯片架构的异构性远超以往的计算市场。个人电脑曾在x86标准和摩托罗拉68x之间竞争,最终形成了今天英特尔和苹果硅片的双头垄断。人工智能的不同用途将需要针对智能、延迟、功耗、训练和推理进行优化的计算。这为专业公司创造了大量空间。一个例子是,ChatGPT 的快速响应模式由 Cerebras 的低延迟硅片驱动。另一个例子是 Fractile(我是其投资者),它与 Anthropic 就其低延迟推理芯片达成了协议。
Compute is becoming a highly segmented market, where chips aren’t a standardized commodity. This differentiated hardware demand will expand the market even if it potentially reduces Nvidia’s relative dominance.
计算正在成为一个高度细分化的市场,芯片不再是标准化商品。这种差异化的硬件需求将扩大市场,即使它可能削弱英伟达的相对主导地位。
Morsels
小片段
Good post from Chad Syverson on AI productivity. Some micro-evidence for improvements, not much at the aggregate level.
Chad Syverson 关于人工智能生产力的好文章。有一些微观层面的改进证据,但在总体层面并不多。
Agglomerations
聚集
Understanding AI and Productivity
理解人工智能与生产力
As part of EIG’s American Worker Project, we are delighted to present this guest post from Chad Syverson, who is the George C. Tiao Distinguished Service Professor of Economics at the University of Chicago Booth School of Business. Click here for a PDF version of this post…
作为 EIG 美国工人项目的一部分,我们很高兴呈现 Chad Syverson 的这篇客座文章,他是芝加哥大学布斯商学院 George C. Tiao 杰出服务经济学教授。点击此处获取本文的 PDF 版本……
Read more
阅读更多
2 days ago · 42 likes · 2 comments · Economic Innovation Group
2天前 · 42个赞 · 2条评论 · 经济创新集团
The least bad place to hide from global catastrophe? Australia.
躲避全球灾难最不糟糕的地方?澳大利亚。
The harness matters as much as the model: SwarmOS pushed GPT-5.6 Sol from 13.3% to 100% on ARC-AGI-3 Public.
控制框架与模型同样重要:SwarmOS 将 GPT-5.6 Sol 在 ARC-AGI-3 Public 上的得分从13.3%提升至100%。
The University of Chicago’s Social Sciences Core is going back to paper, banning most classroom technology to deal with the AI-learning crisis.
芝加哥大学社会科学核心课程将回归纸质教学,禁止大多数课堂技术以应对人工智能学习危机。
Meta considered shrinking some teams by up to 60% to become “AI native.”
Meta 曾考虑将部分团队缩减多达60%,以成为“AI 原生”公司。
Would you take this bet? It pays out if, in any quarter up to and including Q4 2033, US real GDP per capita is at least 15% higher than its previous peak.3
你会接受这个赌注吗?如果在截至2033年第四季度(含)的任何季度,美国实际人均GDP比其先前峰值高出至少15%,则支付赔注。
With one video, you can reconstruct a moving 4D avatar of a person and render it from novel viewpoints, like a video-game character.
通过一段视频,你可以重建一个人的动态4D化身,并从新颖视角渲染它,如同游戏角色一般。
You can now teach an adorable “Pixar-ish” robot new tricks for only $399.
现在,你只需399美元就能教会一个可爱的“皮克斯风格”机器人新技巧。
You can now generate videos in less time than it takes to watch them.
现在,你生成视频的时间比观看视频的时间还要短。
A walk down memery lane:
漫步记忆小巷:
1
1
Reversible processors, like those from Vaire, will not pay the Landauer tax.
可逆处理器,如Vaire公司的产品,将无需支付兰道尔税。
2
2
I run a version of Qwen 3.7-27b on one of our local machines for various tasks.
我在我们本地的一台机器上运行Qwen 3.7-27b版本,用于各种任务。
3
3
Recorded at least four quarters earlier.
至少提前四个季度记录。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力