跳到主内容
@wquguru
精选70Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)行业动态

分析OpenAI与Anthropic在推理RL与预训练上的竞争态势

To be clear this is not a prediction that OpenAI will "win". Anthropic has score…

原文
发到 X

To be clear this is not a prediction that OpenAI will "win". Anthropic has scored a temporary lead by rushing ahead on pretraining and agentic coding. But it seems to me that OAI have a better doctrine on reasoning RL; while Chinese results suggest that scale is overrated for RSI-relevant reasoning capabilities. In combination with OAI's growing quantitative and qualitative compute advantage, they can achieve greater internal R&D velocity, and then they have Astra and whatever is next to catch up on model scale. Even if their science of pretraining is behind Anthropic (which I'm sure it is) and their "10T" is what Anthropic gets with "5T", they only need to maintain positive returns to scale, eat the costs, capitalize on greater knowledge coverage on top of better and more efficient reasoning RL, and then they're back in the lead. The medium-term OpenAI objective, then, is to keep compounding their RL advantage while their next generation pretraining project takes shape.

需要澄清的是,这并非预测 OpenAI 会“赢”。Anthropic 通过在预训练和代理式编程(agentic coding)上加速推进,暂时取得了领先。但在我看来,OpenAI 在推理强化学习(reasoning RL)方面拥有更优的教义;而中国的研究结果表明,规模对于 RSI 相关的推理能力而言被高估了。结合 OpenAI 日益增长的量化与质化算力优势,他们能够实现更快的内部研发速度,随后凭借 Astra 以及后续产品来追赶模型规模。即使他们的预训练科学落后于 Anthropic(我确信确实如此),且他们的“10T”仅相当于 Anthropic 用“5T”能达到的水平,他们也只需维持正向规模回报,消化成本,并在更优且更高效的推理强化学习基础上利用更广泛的知识覆盖优势,便能重新夺回领先地位。因此,中期来看,OpenAI 的目标是在下一代预训练项目成型的过程中,持续累积其强化学习的优势。

Alignment issues (which are inseparable from horrible infra ineptitude) may slow them down, but Anthropic is clearly not perfect here either.

对齐问题(这与糟糕的基础设施无能密不可分)可能会拖慢他们的步伐,但 Anthropic 在此方面显然也不完美。

Bullish on OpenAI.

看多 OpenAI。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近