跳到主内容
@wquguru
精选72Interconnects(RSS)行业动态

开源模型生态的两条未来路径

Teaching Everyone to Fish for Tokens

原文
发到 X

Housekeeping: No voiceover for this post as I’m traveling.

杂务说明:由于我正在旅行,本篇没有配音。

The oldest comparison people try to make is how what’s happening with open models compares to foundational open-source software projects like the Linux operating system. There are fairly clean analogies, but they paint a narrow path forwards for the self-sustaining nature of the open-source model ecosystem, where once Linux got big enough it was going to be self-fulfilling as the best possible tool for many jobs. The open-source language model – i.e. only models that come with a full training recipe, data, code, etc. – is a closer analogue to the open-source operating system. The open weight models you use – those with just model weights and inference code to run them – are closer to specific versions of software that you install in a project built upon them.

人们最常做的比较是,当前开放模型的情况与Linux操作系统等基础开源软件项目相比如何。虽然存在相当清晰的类比,但它们为开源模型生态系统的自持续特性描绘了一条狭窄的前进道路,而一旦Linux足够壮大,它就会在许多任务中成为最佳工具,从而自我实现。开源语言模型——即只有附带完整训练配方、数据、代码等的模型——更接近开源操作系统。你使用的开放权重模型——那些仅包含模型权重和运行它们的推理代码的模型——更接近于你在基于它们构建的项目中安装的特定版本软件。

Share

分享

Model weights are very transient on average, but they still have a long shelf life, as with a lot of heavily used software. It’s why many companies are still using workflows built on Llama 3, despite agentic behaviors taking off years later. The open-source recipe, typified in modern times by the Olmo models I helped build at Ai2, with its predecessors like Pythia from EleutherAI, is a resource intensive process that any company can pick up, modify, and press “run” on to produce a new set of model weights. In the best cases, the community can contribute improvements in data or training code back into the next model too! This is why Nvidia is investing so much in nearly open-source models – for their Nemotron models they release all the data they legally can and the training code, etc. Nvidia wants a world where countless people can build token machines, so intelligence is not monopolized. This is a world with massive demand for inference across many companies, all of which want to buy Nvidia’s offerings.

模型权重平均而言非常短暂,但它们仍然有很长的保质期,就像许多重度使用的软件一样。这就是为什么尽管代理行为在多年后才兴起,许多公司仍在使用基于Llama 3构建的工作流程。开源配方,现代以我在Ai2帮助构建的Olmo模型为代表,其前身如EleutherAI的Pythia,是一个资源密集型过程,任何公司都可以拿起、修改并按下“运行”以生成一组新的模型权重。在最佳情况下,社区还可以将数据或训练代码的改进贡献回下一个模型!这就是为什么Nvidia在近乎开源的模型上投入如此之多——对于他们的Nemotron模型,他们发布了所有合法可发布的数据和训练代码等。Nvidia希望一个无数人能构建令牌机器的世界,这样智能就不会被垄断。这是一个对推理需求巨大的世界,许多公司都希望购买Nvidia的产品。

Open-source AI has a tricky future, as building the best models is extremely capital intensive. The ability to build competitive models has stayed more accessible in industry longer than many would’ve expected. The default expectation for many is that training models is too expensive and the open-source recipe is too far behind, so building a new lab centered on some part of training LLMs will not be tractable.

开源AI的未来充满挑战,因为构建最佳模型极其资本密集。在行业中,构建竞争性模型的能力比许多人预期的更长时间保持可及性。许多人的默认预期是,训练模型过于昂贵,开源配方落后太多,因此围绕训练LLM的某些部分建立新实验室将不可行。

There are two futures from here. First is if “it works” – if the open-source recipe works for Nvidia, they’ll be creating far more demand for their chips (and profits) than it costs to build the models. Right now it’s reported that Nvidia is spending $26 billion on this endeavor. It’s not clear if this will work, or if AI’s capital intensiveness will drive more and more companies out of the training game. We haven’t seen many signs of this starting. In fact, the companies bowing out – like Databricks and 01.ai – seem like anomalies.

从这里出发,有两种未来。第一种是“如果它奏效”——如果开源配方对英伟达有效,他们将为芯片创造远超构建模型成本的需求(和利润)。目前据报道,英伟达为此投入了260亿美元。目前尚不清楚这是否会成功,或者AI的资本密集度是否会迫使越来越多的公司退出训练游戏。我们尚未看到这一趋势开始的许多迹象。事实上,像Databricks和01.ai这样的公司退出,似乎更像是异常现象。

The open-source ecosystem will become increasingly dependent on Nvidia’s financing in the coming years. This is an existential window, where within a few years the profits of this approach need to return to them, or another open model company needs to cultivate platform-like financial feedback loops on their openness. This economic reward needs to be proportional to the profits generated by Anthropic and OpenAI’s APIs to keep pace over decades of language model development. This can be driven by competitiveness on performance or by the AI boom just being so big that the open model training, inference, and fine-tuning companies all have vast quantities of demand.

未来几年,开源生态系统将越来越依赖英伟达的资助。这是一个关乎存亡的窗口期,几年内,这种方法的利润需要回流给他们,或者另一家开源模型公司需要在其开放性上培养类似平台的财务反馈循环。这种经济回报需要与Anthropic和OpenAI的API所产生的利润相称,以在数十年的语言模型开发中保持同步。这可以通过性能上的竞争力或AI繁荣的巨大规模来推动,使得开源模型的训练、推理和微调公司都能拥有庞大的需求。

The second future is if one of these two financially positive paths doesn’t play out, open models will fork to a different development path than the leading closed models – one more focused on efficiency, modifiability, specialization, etc. I put this mentally as my most likely outcome – open models are still incredibly useful, but fill a long-tail ecosystem relative to the closed counterparts that have monopoly ownership stakes in the most valuable areas like knowledge work collaboration, drug discovery, SWE, etc. The long-tail is something like enterprise-specific agents that run on-prem with private data on repetitive business tasks.

第二种未来是,如果这两条财务上积极的道路都没有实现,开源模型将走向与领先的闭源模型不同的发展路径——更侧重于效率、可修改性、专业化等。我在心中将其视为最可能的结果——开源模型仍然极其有用,但相对于在知识工作协作、药物发现、软件工程等最有价值领域拥有垄断所有权的闭源模型,它们填补的是一个长尾生态系统。这个长尾类似于在本地运行、使用私有数据执行重复业务任务的企业特定代理。

Part of why I think this open-source training will have a hard time catching on is because training is getting more complex and more abstracted. The current open model ecosystem is buoyed by an explosion in interest in post-training open models. These people take models like DeepSeek V4 Flash, Inkling Small, or GLM 5.X and finetune them for their specific agentic tasks (e.g. in Tinker, the most popular finetuning API today).

我认为开源训练难以流行起来的部分原因是,训练正变得越来越复杂和抽象。当前的开源模型生态系统受到对后训练开源模型兴趣激增的支撑。这些人拿像DeepSeek V4 Flash、Inkling Small或GLM 5.X这样的模型,针对他们的特定代理任务进行微调(例如,在Tinker中,这是目前最流行的微调API)。

Interconnects AI is a reader-supported publication. Consider becoming a subscriber.

Interconnects AI是一个读者支持的出版物。考虑成为订阅者。

Over the last few years, post-training largely referred to the whole process of modifying the base model to make it intelligent and usable. There is a shift happening where the ability to train a base model to be a general agentic reasoner is becoming opaque like at-scale pretraining practices from a few years ago. This could go so far as to change the established pretraining, midtraining, post-training lexicon that has been standard for a few years. It could come to be something closer to pretraining, reasoning training, and post-training.

在过去几年里,后训练主要指的是修改基础模型使其变得智能和可用的整个过程。现在正在发生一种转变,即训练基础模型成为通用代理推理者的能力正变得像几年前的大规模预训练实践一样不透明。这甚至可能改变已经确立了几年的预训练、中训练、后训练术语体系。它可能会变得更接近预训练、推理训练和后训练。

As there’s less interest in training the entire model, there’s less interest in investing in open-source AI. These are the only sort of hints we will get, but we cannot do much to fight the economic gravity of these situations. This trend is the next step in the number of open model builders who release base models (the model versions before core reasoning training) continuing to decrease. It goes hand in hand with open model builders experimenting with revenue share licenses for downstream use in products or inference. These are experiments in keeping the financing viable for building near frontier open-weight models – a lot hinges in the near future on how successful they are. These are the people that need to succeed for Nvidia’s demand-growth strategy around open-source to succeed, and last.

随着对训练整个模型的兴趣减少,对投资开源AI的兴趣也在减少。这些只是我们能得到的唯一线索,但我们无法对抗这些情况的经济引力。这一趋势是开放模型构建者发布基础模型(核心推理训练之前的模型版本)数量持续减少的下一步。这与开放模型构建者尝试对下游产品或推理使用采用收入分成许可并行。这些是在保持构建接近前沿开放权重模型的融资可行性方面的实验——近期很多取决于这些实验的成功程度。这些人是Nvidia围绕开源的需求增长战略成功并持续所必需成功的群体。

Along the way we’re still in for a ton of action in open-weight models, as releasing access to intelligence is one of the strongest business strategies available. This additional type of player, who monetizes the AI indirectly, is typified by Meta and other hyperscalers with massive balance sheets. Meta releasing its very-strong Muse Spark 1.2 model as open-weights would severely hamper the revenue growth rate of their competitors in Anthropic and OpenAI who rely on selling tokens. These companies are both commoditizing their complements, but they’re doing it in different ways. Nvidia wants to teach everyone to fish for tokens, so the ecosystem is self-sustaining, but Meta is strategically flooding the zone with tokens.

在此过程中,我们仍将看到开放权重模型的大量活动,因为发布智能访问权是可用的最强商业策略之一。这种额外类型的参与者,通过间接方式将AI变现,以Meta和其他拥有庞大资产负债表的超大规模企业为代表。Meta将其非常强大的Muse Spark 1.2模型作为开放权重发布,将严重削弱其竞争对手Anthropic和OpenAI的收入增长率,后者依赖销售令牌。这些公司都在将其互补品商品化,但方式不同。Nvidia想教会每个人如何钓令牌,使生态系统自给自足,而Meta则在战略性地用令牌淹没整个领域。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近