跳到主内容
精选70Gary Marcus(RSS)模型发布/更新多源精选 ×6

OpenAI Astra:惊艳但被过度吹捧的新模型

OpenAI’s amazing — but vastly oversold — new model Astra

原文

Astra, a new model that OpenAI is testing internally, is amazing. No denying that:

Part I: The Biggest Fallacy

But at the same time, a whole raft of people, some fairly prominent, are running around making a deeply flawed argument about the implications.

See if you can spot the fallacy. Here are three examples among many.

Another tweet went so far as to claim that “the species just crossed a one-way threshold”; Elon Musk took it as evidence that we had reached The Singularity.

Each is making essentially the same error.

§

The fallacy has a name; it’s called the fallacy of composition.

The wiki on it is filled with great examples; here are just the first few.

What’s the manifestation of the fallacy in the current case? Thinking that a system that is great at a certain kind of math problem is great at all math, great at science or even quite possibly great at everything. The “AGI-is-near” community keeps committing the same logical fallacy over and over. Every time there’s an advance, I see the same error.

Here’s how the fallacy works.

1. Someone pretends that all cognition is created equally. (Totally untrue.)

2. Whenever AI achieves success on some form of (fancy) cognition, they want you to believe that success on all forms of AI is imminent.

You don’t have to be a cognitive psychologist to realize that this inference just doesn’t follow. We all know, for example, that expertise in math doesn’t guarantee genius in all domains. Someone who is great at math or physics or programming may struggle with writing or understanding human relationships (and conversely a great writer may be weak at math, etc). Expertise in one domain does not at all guarantee expertise in all or even most domains.

That’s precisely *why* people like Howard Gardner and Robert Sternberg developed multidimensional theories of intelligence, why the SAT tests math separately from verbal, etc.

Astra appears to be —we still haven’t seen the methodology—great at math, or at least some forms of math, but that does not mean that it will avoid hallucinations or solve the reliability problems other GenAI systems have. It doesn’t even mean it will be able to read PDFs reliably. And it doesn’t mean it will be the first generative AI to be able to obey hard rules, either. (Which should terrify you.)

In particular, Astra is obviously excellent at some problems; but that doesn’t mean it will be excellent or even competent at problems that are hard to formalize. It doesn’t mean it will be magic. It doesn’t mean it’s AGI or ASI or any of that.

It’s *very* impressive. But I see no reason whatsoever to think Astra is AGI let alone ASI. If it can score even a 5/10 on my 2024 bet with Miles Brundage I will be surprised.

Part II. There is an important, principled reason to think that success on math is a special case which will not generalize as much people might hope.

I wouldn’t go on about this fallacy at such length if (a) it wasn’t wildly common and (b) I didn’t think that there was very good reason to think that the fallacy applied here.

It’s not an accident that math is where these models are shining. Math lends itself to two things: verification (using symbolic tools), and massive amounts of cheaply produced synthetic data where you can guarantee that the answers are correct. The same applies to coding — but it is not true in general. You can generate as many math facts as you want; you can’t simulate the open-ended world. You can verify math; you can’t verify a military strategy in the same way.

This is not news. I have been pointing this out at least since January 2025 [first line should have said coding and math], echoing something Ernie Davis and I said about Go (and why AlphaGo would not be a panacea) in 2019:

OpenAI can do what Astra does in math because math allows for external tools to do verification and to create synthetic data.

That doesn’t make Astra less impressive, but as we learned a decade ago from IBM’s shambolic and ultimately failed attempt to turn Jeopardy-winning Watson into a cancer-fighting machine, success in one domain does not guarantee success in all.

The kind of math OpenAI is dealing with here is radically different from many (perhaps most) real-world problems in which verification neither guarantees solutions nor allows one to produce infinite training data effectively for free.

To expect it to be a universal solvent is to show you don’t understand that basic fact.

Part III: Seven more things to know about Astra

  • Yesterday's tweet and blog were marketing, not science. Neither of those nor the 249-page math article that went with them give any information about how this was accomplished:
  • In an email commenting on the first draft of this essay Ernie Davis added, quite rightly:
  • To properly evaluate the significance of these discoveries we would need some
  • critical pieces of information.
  • First: How many conjectures were attempted? If the OpenAI team picked these 10 conjectures at random from the space of all outstanding significant mathematical conjectures and Astra solved all 10, that would be amazing; but that seems altogether unlikely. If they cherry-picked 50 conjectures that they thought Astra would succeed on and it succeeded on 10, that’s still amazing, but significantly less so. If they ran Astra on all 1000 or so open Erdos conjectures and on 10,000 other open conjectures, then its failure to solve those is significant information on its limits as a mathematical reasoner.
  • Second: OpenAI brags loudly that this was done for a total computation cost of $2000. It seems a safe bet that this includes only the conjectures where Astra succeeded, not the ones where it failed. But what was the cost in terms of the salaries of the highly-paid mathematicians and computer scientists who worked on this project? I’d be astonished if it was less than $20,000 and would not be surprised if it was upward of $200,000.
  • We don’t know this, and the way things are going, we may never know this. We still have zero information about the OpenAI system that achieved gold-medal performance at the 2025 International Mathematical Olympiad, and very little information about the Google system.
  • Astra is good at some math but it is very doubtful that math is “solved” (as a lot of people on X seemed to believe). It seems to be good at certain kinds of math
  • But quite possibly not all. Here are some examples from Eric Weinstein on X about some kinds of math it might be less good at
  • Ernie Davis also adds “The problem of autoformalization --- turning mathematics written by a human mathematician into a strictly logical form -- does not seem close to being solved, though certainly progress has been made. For example: For the last two years, mathematician Kevin Buzzard has been leading a project to formalize Andrew Wiles' proof of Fermat's Last Theorem in Lean. It is quite safe to say that no AI system is capable of carrying out that monumental task. Presumably if they could do it unassisted, they would scoop him, and if they could cut down his work from years to weeks, he would be using them.”
  • Consistent with my general prediction that facility in math doesn’t guarantee universality, it appears that Astra’s proofwriting is not on par with the proofs themselves (echoing something Ernie Davis and I observed here a year ago)” [click through for full tweet]
  • I seriously doubt that Astra will magically solve all the things that don’t work reliably in current models. Will it be able to reliably extract numbers from arbitrary PDFs? Doubt it. Will it solve Sabine Hossenfelder’s desire to have AI write scripts [click the tweet for more details] for her YouTube series? Doubt it.
  • Astra is not (as Musk suggested) the Singularity. Enough said. (But see my recent essay on the topic; nothing has really changed).

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近