OpenAI内部模型一次性发布700多篇数学证明
New Math from OpenAI
这是AI在基础科学领域能力的范式级跃迁,直接证明了前沿模型具备独立解决顶级数学难题的能力,值得所有关注AGI进展的研究者重点关注。
It is kind of a huge deal. OpenAI dumped a broad range of huge new mathematical results produced by an internal frontier model. They just put it all on GitHub.
这简直是个天大的事。OpenAI 发布了一个内部前沿模型产生的一系列重大新数学成果。他们只是把这些内容全部放到了 GitHub 上。
This included 90 of the top 500 open problems in all of math, as per Proof of Atlas. In total there were 722 manuscripts (now 719 after three withdraws from the cluster that did not have Lean proofs) organized into 372 families.
根据 Proof of Atlas 的说法,其中包括了所有数学领域中前 500 个未解问题的 90 个。总共有 722 份手稿(在三个撤回后现为 719 份,因为集群中那些没有 Lean 证明的已被撤回),被组织成 372 个家族。
This was the result of a single model, presumably the same one that produced the Navier-Stokes proof (as per their link back to that post), mostly on a single prompt (quasi-RH was one of the few exceptions), working an average of three hours’ worth of compute per solution found, after being asked to try its luck at about 4,000 problems. The prompt included lines like ‘Even if the problem is “open,” the intention is that you should resolve it and present a full solution.’ OpenAI was trying a lot less than maximally hard.
这是单个模型的结果,很可能是同一个产生了纳维-斯托克斯方程证明的模型(根据其链接回该帖子的信息),主要基于一个提示词(准黎曼猜想是少数例外之一),每个找到的解决方案平均消耗三小时的计算资源,在被要求尝试约 4,000 个问题之后。提示词中包含类似“即使问题是‘开放的’,你的意图应该是解决它并呈现完整的解决方案。”的内容。OpenAI 的努力程度远未达到最大化。
Levant and others called October 6, 2026, ‘obviously the most significant moment in mathematical history.’
Levant 等人称 2026 年 10 月 6 日为“显然是数学史上最重要的时刻”。
Table of Contents
目录
- What Did We Prove?
- The Mathocalypse.
- The World Does Not Understand.
- For Now You Can Still Do Math.
- You Will Need To Find A New Problem.
- Verification or Evaluation Is Not Always Easier Than Generation.
- Cracking the Code.
- Never Change.
- The Mathematicians Are Not Okay.
- The Situation Turns Ugly.
- The Advisory Group Responds.
- Another Mathematical Group Responds.
- A Different Approach.
- Three Withdraws.
- Leaning Into Lean.
- 我们证明了什么?
- 数学末日。
- 世界并不理解。
- 目前你仍然可以做数学。
- 你需要找到一个新问题。
- 验证或评估并不总是比生成更容易。
- 破解代码。
- 永远不变。
- 数学家们状况不佳。
- 局势变得严峻。
- 咨询小组作出回应。
- 另一个数学小组作出回应。
- 不同的方法。
- 三次撤回。
- 深入 Lean。
What Did We Prove?
我们证明了什么?
Here is a thread of common sense graphical explanations (made by AI of course) of top problems. They look like this:
这里有一系列关于顶级问题的常识性图形解释(当然由 AI 制作)。它们看起来像这样:
Here are the text summaries:
以下是文本摘要:
First, the headline is the major breakthrough on Riemann. It’s not a solution to the hypothesis, but it’s a huge tightening on the bounds.
首先,头条新闻是关于黎曼猜想的重大突破。它并不是对猜想的解答,而是对界限的巨大收紧。
Matrix multiplication efficiency at 2.25 may end up having the most real world impact.
2.25 的矩阵乘法效率最终可能会产生最真实世界的影响。
Maybe, but we haven’t yet found a scenario where this is faster in practice.
也许吧,但我们尚未找到在实践中更快的场景。
It’s not quite a practically usable result yet, but matrix multiplication is the majority of the worlds compute. This could open a whole new approach.
这还不是一个真正可用的结果,但矩阵乘法占据了全球计算的大部分。这可能开启一种全新的方法。
Integer multiplication is probably the most shocking result. Similar to matrix multiplication it’s a “asymptotic reduction”. Also like the matrix result, it’s not practical.
整数乘法可能是最令人震惊的结果。类似于矩阵乘法,它是一种“渐近简化”。与矩阵结果一样,它也不具备实用性。
But very few people would have predicted the previous barrier could be broken. Even slightly
但很少有人会预测到之前的壁垒能被打破。哪怕只是稍微突破一点。
The Pi result isn’t necessarily the most important result here, but it may be the most fun. And probably accessible enough that even your kids can get it. Basically just how irrational is Pi.
圆周率(Pi)的结果不一定是这里最重要的,但它可能是最有趣的。而且可能足够通俗易懂,连你的孩子都能理解。基本上就是关于 Pi 有多无理的问题。
Unique games is also a pretty huge results, because like Riemann it’s adjacent to one of the most important problems in math: P vs NP.
唯一游戏(Unique games)也是一个相当大的成果,因为就像黎曼猜想一样,它与数学中最重要的问题之一 P vs NP 相邻。
To be clear it’s not about P = NP itself. But does reveal new things about the fundamental limits to NP hard problems.
要明确的是,这并不是关于 P = NP 本身。但它揭示了关于 NP 困难问题基本限制的新见解。
Technically two problems, but I put them together because of similarities, in proof and implications.
从技术上讲是两个问题,但我把它们放在一起,因为它们在证明和推论方面有相似之处。
Hodge and Birch are both pretty abstract. But they’re both so important to algebraic geometry that after Riemann they may he the most significant results in the set.
霍奇猜想(Hodge)和伯奇-斯温纳顿-戴尔猜想(Birch and Swinnerton-Dyer)都相当抽象。但它们对代数几何都非常重要,因此在黎曼猜想之后,它们可能是该组中最具意义的成果。
Hilberts 10th is another computer science problem. But unlike P vs NP it isn’t even about problems that are computationally hard to solve, it’s about whether a type of problem even has a computable solution.
希尔伯特第十问题是另一个计算机科学问题。但与 P vs NP 不同,它甚至不是关于在计算上难以解决的问题,而是关于某类问题是否根本存在可计算的解。
Last but not least is Hadwiger on graph coloring. It also ranks there for most shocking result of the full set.
最后但同样重要的是哈德格维希(Hadwiger)关于图着色的研究。它在整个集合中也以最具震撼力的成果排名靠前。
It basically overturns something that we thought was one of the most fundamental relationships in graphs
它基本上推翻了我们认为的图中最基本的关系之一。
Here are some analyses of the results in algebraic number theory.
以下是一些关于代数数论结果的深入分析。
What is all of that good for? Ole Lehmann’s AI mentions applications for fusion research, portable body scanners, matching systems, tissue scans, quantum sensors and safety checks on self-driving cars and robots, among other things.
这一切有什么用呢?Ole Lehmann 的 AI 提到了其在聚变研究、便携式人体扫描仪、匹配系统、组织扫描、量子传感器以及自动驾驶汽车和机器人的安全检查等方面的应用,以及其他用途。
The Mathocalypse
数学末日(The Mathocalypse)
Alex Kontorovich: Quasi-RH?!?!???! Are you kidding me? If a human did this, it would be an instant Fields Medal, no questions asked.
Alex Kontorovich:准黎曼假设(Quasi-RH)??????你在开玩笑吗?如果是人类做到的,这将毫无争议地立即获得菲尔兹奖。
RH says zeta has no zeros in Re(s)>1/2. The best we had until a second ago was a region that got thinner and thinner the higher up the imaginary axis you go. I thought maybe they’d fatten that up a bit, that’d be a massive breakthrough. But no. They got a zero free strip!!!! Insane
黎曼假设(RH)指出 zeta 函数在 Re(s)>1/2 区域没有零点。直到刚才为止,我们最好的结果是随着虚轴向上延伸,无零点区域变得越来越窄。我原以为他们可能会让那个区域宽一点,那将是一个巨大的突破。但没有。他们得到了一个无零点带!!!!太疯狂了
Yeah, no Siegel zeros either. So I guess two Fields medals…
是的,也没有西格尔零点(Siegel zeros)。所以我猜是两枚菲尔兹奖……
Quasi-RH is not the full Riemann hypothesis, but it is sufficient for many purposes, such as computing square roots modulo a prime quickly without coin flipping, or getting a much better estimate of the number of primes below a given number, and a sibling paper gets to the core of Artin’s 1927 primitive root conjecture.
准黎曼猜想(Quasi-RH)并非完整的黎曼猜想,但它足以用于许多目的,例如在不使用随机抛硬币的情况下快速计算模素数的平方根,或对给定数值以下的素数数量获得更精确的估计;而一篇姊妹论文则深入探讨了阿廷1927年原始根猜想的本质。
Steven Strogatz: Many staggering results here. But this is a particularly amazing one: the exponent for matrix multiplication is no more than 2.25. The previous world record had been something like 2.37. This leap in progress is like Bob Beamon’s long jump.
史蒂文·斯特罗加茨:这里有许多令人震惊的结果。但这一条尤为惊人:矩阵乘法的指数不超过2.25。此前的世界纪录约为2.37。这一进展飞跃堪比鲍勃·比蒙的跳远壮举。
Isaac Kim: This contains a shocking list of problems in quantum information, many-body physics and quantum computing, the field that are dear to my heart. There are too many, but let me pick the following.
艾萨克·金:这份清单令人震惊,涵盖了量子信息、多体物理和量子计算领域中的诸多问题——这些正是我珍视的领域。问题太多,但让我挑选以下几项。
1. Proof of area law in 2D.
1. 二维面积律的证明。
2. Spin-one Haldane gap
2. 自旋-1霍尔丹能隙
3. Parity is not in QAC^0
3. 奇偶性不属于QAC^0
4. Constant-error Aaronson-Kuperberg conjecture
4. 常数误差的Aaronson-Kuperberg猜想
5. Unitary VOAs generating conformal nets.
5. 酉顶点算子代数生成共形网。
Things are changing fast. I cannot even imagine what will happen over the next few months, let alone a year.
变化正在加速。我甚至无法想象未来几个月会发生什么,更不用说一年了。
will depue: i asked GPT 6 Pro and Fable 5.1 to rank all discoveries in the last three years
will depue:我问过GPT 6 Pro和Fable 5.1对过去三年的所有发现进行排名
for Human discovered
人类发现的
for AI discovered (before October 6th)
AI发现的(截至10月6日之前)
for AI discovered from OpenAI/math repo
从OpenAI/math仓库中AI发现的
81% of them have been released today. wtf
其中81%今天已经发布。真是疯了
Fable 5.1’s top 100 were 59% today’s list, 87% AI, full list at the link.
Fable 5.1的前100名中有59%出现在今天的列表中,其中87%为AI发现,完整列表见链接。
Sauers: My agents (Opus 5.5, Astra, and 5.6 Sol) already solved some of these beforehand (on GitHub). I wonder what % of these actually require their internal model?
Sauers:我的代理模型(Opus 5.5、Astra和5.6 Sol)此前已经解决了其中一些问题(在GitHub上)。我想知道这些结果中有多少比例真正依赖于它们的内部模型?
The World Does Not Understand
世界尚未理解
The AIs all say when asked about all these math proofs as a hypothetical that this would be a huge deal, and should be front page news, historic beyond any reasonable comparison, the biggest day in mathematics. Some think it is impossible.
当被问及这些数学证明作为假设情境时,所有AI都表示这将是一件大事,应成为头版新闻,其历史意义远超任何合理比较,堪称数学界最大的一天。有些AI认为这不可能实现。
It turns out this was not front page news. Most people did not hear about it. It should have been, but no one in the news business knows what it means, or thinks people would care.
事实证明,这并未登上头版新闻。大多数人对此一无所知。它本应如此,但新闻媒体从业者要么不知其含义,要么认为公众不会关心。
Kevin A. Bryan: Ok, finished going through the OAI math list. It is bonkers and should be front page news around the world if folks understood what this meant. But if I am running strategy at a lab, #1 priority is “do this for medicine, oncology, battery efficiency, etc as fast as possible”.
凯文·A·布莱恩:好的,我已经浏览完OAI数学列表。这简直疯狂,如果人们理解其含义,理应成为全球头版新闻。但如果我是某实验室的战略负责人,首要任务将是‘尽可能快地将其应用于医学、肿瘤学、电池效率等领域’。
Joshua Gans starts off his coverage with the actual front page, in order to show us what things The New York Times thought were more important than solving 90 of the top 500 open math problems at once.
乔舒亚·甘斯在其报道开头展示了真正的头版内容,以此向我们表明《纽约时报》认为哪些事情比一次性解决500个顶级开放数学问题中的90个更为重要。
Seth Burn: Sometimes front page news doesn’t initially reach the front page. The best example is Sputnik. It launched 10/4/57, but the Soviets didn’t think it was a huge deal [and it got a few paragraphs of a right-side column].
Seth Burn:有时头版新闻最初并未登上头版。最好的例子是斯普特尼克号(Sputnik)。它于1957年10月4日发射,但苏联人并不认为这是件大事[因此仅在右侧专栏占了几段篇幅]。
It was only after America freaked out that it [became a banner headline] in Russia on 10/6/57.
直到美国方面大惊失色后,它才在1957年10月6日成为俄罗斯的头版头条。
It was a remarkably slow news day otherwise. Much of this is not exactly breaking. Yet no math is to be seen, even in the summaries at the bottom, which even include a generic AI thinkpiece.
除此之外,这是一个新闻极其平淡的日子。其中大部分内容算不上突发新闻。然而,即便是在底部的摘要中,也看不到任何数学相关内容,这些摘要甚至包含了一篇通用的AI评论文章。
Joshua Gans: What’s missing is this: 722 mathematics papers written by OpenAI. It isn’t an overstatement to say that this is probably the biggest day of scientific advancement in history. It is hard to evaluate, but the discussion that I have been seeing is that many of these are among the hardest and most significant results in mathematics. I suspect October 6th, 2026, will go down as some form of Judgment Day for AI in mathematics, but it portends so much more.
Joshua Gans:缺失的是这一点:OpenAI撰写的722篇数学论文。说这可能是人类历史上科学进步最大的一天也不为过。这很难评估,但我看到的讨论表明,其中许多论文属于数学中最困难且最重要的成果。我怀疑2026年10月6日将成为AI在数学领域的某种“审判日”,但它预示的远不止于此。
The counterargument is that there might be a lot of days like this:
反方观点是,类似这样的日子可能还有很多:
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力