跳到主内容
@wquguru
精选88The Zvi(RSS)模型发布/更新

OpenAI新模型8天解千禧难题,引发学术诚信争议

Brand New AI Solves a Millennium Prize

原文
发到 X
推荐理由

这不仅是模型能力的里程碑式展示,更揭示了AI介入基础科学后引发的学术伦理与署名权冲突,对科研圈影响深远,值得密切关注后续进展。

The first Millennium Prize, Navier-Stokes, has fallen to AI.

第一个千禧年大奖难题——纳维-斯托克斯方程,已被人工智能攻克。

A deeply unfortunate situation has arisen involving what should have been some combination of a positive story about new progress in AI-assisted mathematical research and yet another opportunity to freak out about rapid AI progress.

出现了一个极其不幸的局面,这原本可以是一个关于AI辅助数学研究取得新进展的积极故事,以及又一次让人对AI快速进步感到惊恐的机会。

Or, as we call it around here, Tuesday.

或者,就像我们这里常说的:星期二。

The Real Story Is The New Model That Is Better Than Astra

真正的故事是那个比Astra更强大的新模型

Keep your eyes on the prize. There are three stories here.

紧盯目标。这里有三个故事。

The first story is much more important than the second story, which in turn is much more important than the third story.

第一个故事远比第二个故事重要,而第二个故事又远比第三个故事重要。

  • OpenAI’s next model took a week to get a generation ahead of Astra, and they are telling us this because everyone is totally freaked out about what is happening.
  • Or, in official language: ‘We believe it is important to inform the world about the pace of AI progress and what to expect from upcoming models’ and that we ‘may require more deliberate choices about the pace of progress.’
  • This new AI has, eight days after it started training, solved Navier-Stokes.
  • A bunch of drama over who gets the credit for math involving Navier-Stokes.
  • OpenAI的下一个模型花了一周时间才在代际上超越Astra,他们之所以告诉我们这些,是因为每个人都对正在发生的事情感到极度恐慌。
  • 或者用官方语言来说:“我们认为有必要告知世界AI进步的节奏以及未来模型可能带来的预期”,并且我们“可能需要更审慎地选择进步的节奏。”
  • 这个新的AI在开始训练八天后,就解决了纳维-斯托克斯方程问题。
  • 围绕谁该为涉及纳维-斯托克斯方程的数学成果获得荣誉而引发的一堆戏剧性事件。

Jeffrey Ladish: I haven’t looked into the human drama around the Navier–Stokes problem but sorry give me a minute because HOLY SHIT AI JUST SOLVED A MILLENNIUM PROBLEM.

杰弗里·拉迪什(Jeffrey Ladish):我还没深入调查围绕纳维-斯托克斯问题的人类戏剧,但抱歉让我缓一下,因为天哪,AI刚刚解决了一个千禧年大奖难题。

I cannot emphasize the top story enough.

我怎么强调头条新闻的重要性都不为过。

I am still going to tell all three stories, but again: Eyes on the prize.

我仍然会讲述这三个故事,但再次强调:紧盯目标。

Setting the Stage

铺垫背景

The story on this particular Tuesday begins in the morning.

这个特定星期二的故事从早晨开始。

Tristan Buckmaster and Levent Alpoge had worked for a year and offer us a series of remarkable results: Finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3D incompressible Euler. They also believe they have a blowup for hypo-dissipative Navier-Stokes, but the Lean verification of that result is not finished.

特里斯坦·巴克马斯特(Tristan Buckmaster)和莱文特·阿尔波格(Levent Alpoge)工作了一年,向我们提供了一系列非凡的成果:不可压缩多孔介质、布西内斯克方程(Boussinesq)以及三维不可压缩欧拉方程在光滑外力作用下的有限时间爆破。他们还认为自己在假设耗散纳维-斯托克斯方程方面取得了爆破结果,但该结果的Lean验证尚未完成。

Tristan Buckmaster: The program this fits into was not started by us nor was it proposed by a Large Language Model. The credit for the basic idea of this program goes to Diego Cordoba and Luis Martınez-Zoroa, who for several years have been

特里斯坦·巴克马斯特:这个项目并非由我们发起,也不是由大型语言模型提出的。该项目基本想法的功劳归于迭戈·科尔多瓦(Diego Cordoba)和路易斯·马丁内斯-佐罗阿(Luis Martinez-Zoroa),他们在过去几年里一直

exploring the construction of forced blow ups. We took their work as a starting

探索受迫爆破的构造。我们以他们的工作为起点,

point, using Large Language Models to push their program to completion.

利用大型语言模型推动他们的项目完成。

Concretely, what Levent and I did was to take the Cordoba and Martınez-Zoroa program, which achieved blowup results with rough forcing, and, with a

具体来说,莱文特和我所做的是,将科尔多瓦和马丁内斯-佐罗阿的项目(该项目通过粗糙外力实现了爆破结果)与一个

great deal of help from LLMs, push it to smooth forcing and to the incompressible Euler equations. The ideas making this line of attack possible are due to

借助大语言模型的大量帮助,将其推向光滑强制解以及不可压缩欧拉方程。使这一攻击路线成为可行的思想归功于

Cordoba and Martınez-Zoroa.

Cordoba 和 Martınez-Zoroa。

Let me make plain what I have said to colleagues in private: in view of this body of work, I believe Luis Martınez-Zoroa deserves a Fields Medal.

让我私下里对同事们明确表达我的观点:鉴于这一系列工作,我认为 Luis Martınez-Zoroa 值得获得菲尔兹奖。

So far this is great. As he notes requires rethinking about how all of advanced math will function going forward. Terence Tao offers commentary on the underlying results.

到目前为止这非常棒。正如他所指出的,这需要重新思考所有高等数学在未来将如何运作。Terence Tao 对基础结果发表了评论。

The part that is not so great is where they felt forced to publish early, without the opportunity to spend the weeks necessary to make the proofs what passes among mathematicians as readable. Buckmaster outright apologizes for the way the reports look, comparing the Euler writeup in particular to AI slop.

不太好的部分是他们感到被迫提前发表,没有机会花费必要的几周时间,以使证明达到数学家们认为的可读标准。Buckmaster 直接为报告的外观道歉,特别将欧拉方程的撰写比作 AI 生成的垃圾内容。

What happened?

发生了什么?

I Heard a Rumor

我听到一个谣言

On September 1, OpenAI heard a (false) rumor that two Millennium Problems had been resolved. In response to the possibility that someone else might make the world a better place, OpenAI spent millions of dollars in inference to explore all the unsolved Millennium Problems, using an internal model stronger than Astra, and cracked the whole of Navier-Stokes in 88 hours, plus another 17 for Astra to do the Lean formalization and verification.

9月1日,OpenAI 听到了一个(错误的)谣言,称两个千禧年大奖难题已被解决。为了应对其他人可能让世界变得更美好的可能性,OpenAI 投入数百万美元进行推理,以探索所有未解决的千禧年大奖难题,使用了比 Astra 更强的内部模型,并在88小时内破解了整个纳维-斯托克斯问题,另外又花了17小时让 Astra 完成 Lean 的形式化和验证。

You can view this as ‘OpenAI got curious to see what their new baby could do’ or you can view this as ‘they spent millions trying to scoop what they thought was an Anthropic project.’ My money is on a little from column A, a little from column B.

你可以将其视为‘OpenAI 好奇地想看看他们的新宝贝能做什么’,也可以将其视为‘他们花费数百万美元试图抢先发现他们认为属于 Anthropic 的项目’。我认为两者兼有。

All it took was the rumor.

这一切只需一个谣言。

Our Price Cheap

我们的价格低廉

OpenAI: Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens. In the process of resolving the Navier–Stokes problem, the agents sent 2.7 million messages and used approximately 130 billion output tokens.

OpenAI:在所有尝试的问题中,智能体发送了490万条消息,并使用了约3000亿个输出标记。在解决纳维-斯托克斯问题的过程中,智能体发送了270万条消息,并使用了大约1300亿个输出标记。

Depending on what costs you count, this would have cost a regular customer on the order of $22 million, or several million for internal marginal costs. A small price if it works, but also it is crazy how many people don’t think two steps ahead.

根据你计算的成本不同,这对普通客户来说成本约为2200万美元,或者内部边际成本为几百万美元。如果有效,这是一个小价钱,但也疯狂的是,竟然有那么多人不会多想两步。

roon (OpenAI): like all other technologies your equivalent agent swarm will cost a buck fifty in like a year. all this stuff will mean an unprecedented Enlightenment no matter the costs today

roon (OpenAI):像所有其他技术一样,你的等效智能体群集在一年内成本将仅为1.5美元。无论今天的成本如何,所有这些都将带来前所未有的启蒙时代。

I mean, no, maybe $150k or if you’re lucky $15k, but the point stands.

我的意思是,不,也许是15万美元,或者如果你运气好是1.5万美元,但道理是一样的。

Sam Altman (CEO OpenAI): ugh AI is such a bubble, i heard they are selling tokens at a loss, did they know this was only worth $1 million?

Sam Altman (OpenAI CEO):唉,AI 真是个泡沫,我听说他们在亏本出售代币,他们知道这仅价值100万美元吗?

A Millennium Prize result is worth vastly more than a million dollars. OpenAI does not intend to claim the prize money. This was never about prize money, for anyone. It was always about the credit.

千禧年大奖难题的奖金价值远超一百万美元。OpenAI 无意申领这笔奖金。这从来就不是为了奖金,对任何人而言都不是。它始终关乎的是署名权与荣誉。

An Accusation Is Made

一项指控被提出

After hearing internet rumors, Buckmaster reached out to OpenAI, and according to Buckmaster OpenAI’s Sebastien Bubeck said that an internal OpenAI model had produced a proof of finite time blowup for the forced Navier-Stokes equations over the past few days. Buckmaster claims that over the course of two calls, it became clear he had initially been misled and that an entire team had been working on the problem, using an ‘insane’ amount of compute.

在听到网络传言后,Buckmaster 联系了 OpenAI。据 Buckmaster 所述,OpenAI 的 Sebastien Bubeck 表示,过去几天里一个 OpenAI 内部模型已经生成了关于受迫 Navier-Stokes 方程有限时间爆破(finite time blowup)的证明。Buckmaster 声称,经过两次通话,他逐渐意识到自己最初受到了误导,而且整个团队一直在解决这个问题,使用了“疯狂”算力的投入。

Which we now know was 130 or 300 billion output tokens, depending on what counts.

我们现在知道,这相当于 1300 亿或 3000 亿个输出 token,具体取决于如何计算。

If this part is true, it would be quite bad.

如果这部分属实,那情况会相当糟糕。

Tristan Buckmaster: Two proposals were offered to me.

Tristan Buckmaster:有人向我提出了两项提议。

The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day.

第一项是我们发布我们的 Euler 方程结果,然后 OpenAI 在第二天发布他们的 Navier-Stokes 方程结果。

The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us,

第二项是,在我们发布 Euler 结果后,由我单独撰写一篇论文来呈现 Navier-Stokes 的结果,并承认该结果是由一个 OpenAI 内部模型解决的。Sebastien 两次主张将 Levent 从作者名单中移除,并表示如果仅仅是因为 Levent 在 Anthropic 工作这件事如此令人恼火且并非事实的话,一切都会很简单。此外还提到,如果 OpenAI 在我们之后发布结果,

they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers.

他们会说我们配得上克莱研究奖(Clay Prize),并称我们是“最接近解决该问题的人类”。我拒绝了这两项提议。

I said that if OpenAI released its result in the way proposed I would go

我说,如果 OpenAI 按照所提议的方式发布其结果,我将

public with what happened. The reply was, “Why would you ruin your career?”

公开此事。得到的回复是:“你为什么要毁掉你的职业生涯?”

I replied that I am an academic, and asked why he thought going public would

我回答说我是学术界人士,并问他为什么认为公开此事会

ruin my career. The reply was, “If you don’t want me to be nice, then I don’t

毁掉我的职业生涯。回复是:“如果你不想让我客气,那我就不必客气。”

have to be nice.”

have to be nice.”

… I would like to be clear about what I am not claiming. I have not seen

……我想澄清一下我没有声称的内容。我没有见过

OpenAI’s proof. I do not know what their model did, or how. I do not know

OpenAI 的证明。我不知道他们的模型做了什么,或者是怎么做的。我不知道

whether our data was used. I am not accusing anyone of anything. I am stating

是否使用了我们的数据。我没有指控任何人做任何事。我只是陈述

what I was told, when, and what was proposed to me. I am stating it because

我被告知了什么、何时被告知以及对我提出了什么提议。我之所以陈述这些,是因为

the alternative is to let a sequence of announcements say something I know to

如果不这样做,就任由一系列公告说出一些我知道

be false.

是错误的东西。

Or here’s Levent Alpoge telling his side of the story, and Sebastien responding:

或者这里是 Levent Alpoge 讲述他的版本的故事,以及 Sebastien 的回应:

levent: “we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”

levent:“我们不能排除源自他们使用我们产品的去标识化数据有助于改进我们模型的可能性。”

i mean props to them for straight coming clean.

我的意思是,佩服他们直接坦白。

(so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan)

(到目前为止,这个证明看起来更像是我们之前另一个欧拉爆破证明的路子,基于其假设的命名,我们开了些非常愚蠢的双关语,比如“smooth criminale”,不像那个好得多的“ideal fluids explode”,Tristan)

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件