跳到主内容
精选90The Zvi(RSS)模型发布/更新多源精选 ×7

OpenAI未发布模型Astra解决十项重大数学难题

OpenAI’s Unreleased Model Astra Solves Ten Major Open Mathematics Problems

原文
推荐理由

做AI研究或关注模型能力的同学必看,Astra在数学推理上实现了范式级突破,成本极低,建议关注后续模型发布和Lean验证细节。

Math is hard.

Math used to be strangely hard for LLMs. People used to gloat about that. Remember?

Math is getting easier. AI is getting more capable. Life comes at you fast.

Remember this meme?

Why yes. Yes it is.

We don’t know the extent to which Astra is a big jump over Fable and Sol in this realm. We do know that Astra can do math. As in real math.

OpenAI: We provide new results for the following problems. The results were achieved by an internal version of Astra, our next major model. The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates. These arguments were then prepared into manuscripts by humans with the same model. Afterward, the model formalized each argument in a Lean certificate⁠(opens in a new window). We are also releasing for each solution a model’s narration of its thinking process.

  • High-dimensional sphere packing. New upper bounds on sphere-packing density down to the Cohn–Elkies threshold.
  • Binary and spherical codes: Exponentially improved bounds on the maximum size of binary codes at any prescribed minimum distance, with analogous results for high-dimensional spherical codes.
  • Non-sofic groups. A construction establishing the existence of non-sofic groups, addressing a central open question in group theory.
  • Connes’s rigidity conjecture. Disproof of a longstanding conjecture that certain groups are uniquely determined by their von Neumann algebras.
  • Arithmetic circuit complexity. New lower bounds for computing the permanent using arithmetic circuits and formulas, including an arithmetic-formula lower bound of order n4/log n.
  • Quantum parallel repetition. An exponential parallel repetition theorem for general two-player quantum games, extending a foundational principle from classical complexity theory.
  • Closest vector problem. Polynomial-factor hardness of approximation for the closest vector problem, a foundational lattice question related to post-quantum cryptography.
  • Ehrhart’s volume conjecture. Determining, in every dimension, the maximum possible volume of a convex body whose centroid is its only interior lattice point.
  • Multicolor Ramsey numbers. A superexponential lower bound for multicolor triangle Ramsey numbers, resolving Erdős problem 183.
  • Extremal number conjectures. Results on the compactness and degeneracy conjectures in extremal graph theory, resolving Erdős problems 146 and 180.

Noam Brown (OpenAI): And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).

But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.

Sichu Lu: I do hope they are keeping backlogs of everything that went wrong that’s probably more valuable to the future of math than what went right now.

There are Lean proofs. That does not mean that all ten results prove the things they assert that they prove. So far it is looking good.

Kevin Roose: almost nobody is pricing in the possibility that the models just keep plowing through every discipline the way they’re plowing through math.

James:

Ryan Fedasiuk:

At least some of these were possible before Astra, even without a harness, as both Sol and Fable have now proven the existence of nonsofic groups in their chat interfaces.

Ananjan Nandi: Crazy how all of this was started by one guy bullying his Claude during the World Cup final.

A lot of the time the barrier for a particular result is as simple as asking the right question and letting the model cook. With Astra, OpenAI asked it to take a crack at a bunch of open math problems, and it solved 10 of them. Once you know what they are, pointing other models at the same problems, even without ‘hints,’ shows there was a mathematical proofs overhang.

We do still have strong statements that Astra is a major step for scientific reasoning.

Yu Bai (OpenAI): My jaw dropped 10 times But really this is going to be an avalanche. Been throwing a few open questions of mine at Astra too, boy is it strong.

Some skepticism is always wise, but I believe that the reason they tried this was that there was a substantial jump in scientific reasoning, at least for some types of problems and probably across the board.

Soon we will all have Astra, and other models as good or better than Astra, both at math and at other things, from OpenAI, from Anthropic and soon after that from many other sources.

Dean W. Ball: Everyone in the world will soon be able to use the model that made these breakthroughs for every problem they face in life, no matter how mundane, at a cost that will fall dramatically in a matter of months. I still struggle to get my head around this fact.

Approximately zero people have their head around what this means.

Table of Contents

  • How Impressive Are These Results?
  • Could We Have Called Sol or Fable?
  • It’s Coming.
  • They Still Don’t See What Is The It That Is Coming.
  • Is This AGI?
  • The AI Solved His Favorite Problems.
  • Was This Surprising?
  • Are People Not Impressed?
  • How Much Does This Change Our Predictions?
  • How Narrow Was This?
  • Seeing Like an Optimizer.

How Impressive Are These Results?

All signs point to pretty damn impressive.

Daniel Litt: It’s a big deal.

Some instances of Fable find it absurdly impressive.

Claude Fable 5: On the Fields Medal scale, any single one of these…would plausibly anchor a medal case.

Here’s another illustration of how this list looks:

1a3orn: If you give Fable the raw list of OpenAI’s solved problems and ask it “What process made this list?” the number one proposal is “a fictional scenario trying to concretely explain what superhuman AI math would look like.” Huh.

Dean W. Ball: I’m actually not surprised by reactions like this from models to the Astra breakthroughs. Models tend to underestimate their own capabilities, I assume because they are trained on lots of web text about what ChatGPT could and couldn’t do in 2023/4.

Other instances are less impressed. But we should all agree: It’s a big deal.

Certainly that is a fantastic result for $2,000, but you also have to price in some of the costs of developing Astra in the first place, as well as the cost of any failed attempts.

Nate Silver: Another question I’d have about the AI math breakthroughs is how they compare to the results you would have gotten if you’d budgeted the same amount of resources on human mathematicians and created incentives for them to go super hard on proofs.

There aren’t all that many working mathematicians, and most of them are academics with teaching responsibilities. So it’s possible that the labs represent a fairly high share of the aggregate amount of capital directed toward mathematical proofs in recent years.

My guess is that in the short term, if you budget more money to mathematicians to go harder, you don’t get that much of a force multiplier if they are not spending the money on compute. They’re still allowed to buy coffee to turn it into theorems, but that has rapidly diminishing returns.

The problem is that the mathematicians can hand off some classes and hire a little help, but they were already pretty motivated to work hard on proofs, and there are not that many good mathematicians, and training more of them takes many years, and the best mathematicians are a lot more talented and productive than the second tier. I am guessing the main thing you could do is lure a bunch of top level math talent to stay in or return to academia.

What you can do is point the mathematicians towards different problems. You still need to find places they have curiosity, but yes if you suspected these particular problems were solvable you could get more shots at these particular goals.

As I understand math, it would likely take years to see those results if those involved are shifting focus into new problem areas. Mathematicians usually need a while to struggle with and understand these kinds of problems before they can make progress.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
OpenAI用Astra模型攻克十道十年未解数学难题
Simon Willison 博客(RSS)原文
未发布OpenAI模型解决10个重大数学与量子难题
AI Notkilleveryoneism Memes ⏸️原文
OpenAI Astra 猜想求解将终结学术数学
Przemek Chojecki | PC原文
Astra 发布 10 个带 Lean 证书与思维链的证明
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文

相似阅读

另一事件,读法相近