跳到主内容
精选70The Verge AI(RSS)行业动态多源精选 ×2

AI 冲击数学界:顶尖数学家陷入存在危机

Welcome to the AI crisis in math

原文

Today on Decoder, I’m talking with Robert Hart, The Verge’s London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis many lead mathematicians are having about it.

OpenAI just published a set of solutions to longstanding problems in math that went off like a bombshell in the field. It caused a huge debate in the math community, and Rob spent some time talking to some of the most accomplished mathematicians of our time about it.

It’s funny that AI systems are all still pretty bad at elementary school arithmetic, but getting increasingly good at very high-end abstract math. That raises some big questions for the field of advanced math.

If AI can do math of this caliber, does that mean AI labs can transfer those skills to other domains? What good are academic grants and university programs training new generations of human mathematicians to identify new problems as they try to solve existing ones, if frontier models simply answer all the outstanding questions?

What if all this attention around math is just a big marketing exercise for frontier AI labs, which couldn’t care less what happens to one of the oldest and most fundamental academic disciplines there is?

There’s a lot here, and Robert has talked to a lot of people with a lot of views on all of it.

Okay: Verge AI reporter Robert Hart on what AI is doing to math. Here we go.

This interview has been lightly edited for length and clarity.

Robert Hart, you’re our London-based AI reporter here at The Verge. Welcome to Decoder.

Thank you for having me.

I am very excited to talk to you. There’s a lot going on in particular with AI and math that you recently dove into. You spoke to a lot of leading mathematicians about the crisis in mathematics due to AI.

It feels like a lot, and also like there’s a lot yet to know and discover about the interaction of these two things. A full existential crisis, which is pure Decoder bait. Broadly tell us what’s going on.

I think “a lot” sums it up quite well. Basically a bit of an existential crisis within, “what is mathematics? What are mathematicians doing, and what is the role of mathematicians going forward?”

A lot of that has been spurred by a phrase transition in what AI is capable of that has exploded in the last six months to a year. AI went from being very terrible to seemingly genuinely quite good at a professional level in a very short space of time. It’s a lot of what these other fields have been struggling to deal with for the last five years in a very compressed period of time.

I would put that next to software engineering. We’ve been living through the AI crisis in software engineering for some amount of time. But as recently as 2024, even last year, the conventional wisdom was that AI models were particularly bad at math. The famous example is that these models could not count the number of R’s in the word strawberry. Even just counting eluded them.

What has happened to make them better at math? Are they still bad at general arithmetic and they’re good at advanced math, or is it something in between?

They are still truly, truly terrible at some areas of math. I did check, they can do strawberries now. I think someone’s tweaked it.

I think strawberry’s hard-coded. I want to be very clear, my conspiracy theory is that the strawberry thing is hard-coded into all the models.

I think so too. That is a conspiracy I’ll buy into.

But yeah, it’s still terrible at those kinds of things — math, arithmetic, even the days of the week. My boyfriend was saying the other day, “It keeps thinking it’s Wednesday. It’s not Wednesday.” Or time. Elissa Welle for us a few months ago wrote that ChatGPT can’t tell time. Still can’t. That’s not all of math.

So there’s this disconnect. To be good at math, you’ve got to be good at counting, or adding, or multiplying. A lot of it is actually reasoning. If you look at academic math papers, a lot of the time you won’t see numbers, which sums that one up, I think. So they’re still terrible, but they’re now also very good at this other part.

As to why, at some point you reach a critical mass of what these systems can do. We saw it with writing, we’ve seen it with programming. They’re very good at forging connections between different areas, applying old methods in new ways, those kinds of things. It appears that the newer models they’re training have apparently reached that level where it clicks, and now it can do math.

It’s important to say as well that we speak of math as a unitary discipline, especially from the outside. But imagine, say, biology. You’ve got something that would range from literally watching animals and describing behavior all the way through to cellular mechanisms and biochemistry. Math is not a unitary discipline either. AI is really good at some bits. Some bits like counting, it’s still really bad at.

Even on the more abstract levels, mathematicians have floated topology as one area that AI apparently still quite bad at. I can’t verify that, to be honest. It’s beyond my area of expertise. Still, it’s a bit of a mixed bag.

So you’ve described mathematics as a huge field, obviously, with many, many academic areas of interest. There are some parts where the models have gotten quite good. There are other parts, maybe the basic parts that people think of as math, which is simply counting, where they’re still struggling, and then there’s a wide range in the middle.

Is it the wide range in the middle where the existential crisis is, people don’t know what’s going to happen? Or is it at the parts where it’s really good?

A bit of both, which I feel is going to be a running theme through this. No one’s really afraid of it being a mediocre mathematician, but obviously there’s a huge element of what this field does. In terms of the research elements or the cutting edge, as we see with a lot of the results that generate hype, what can it do? There are areas now where it seems to be producing work that is on par with good mathematicians, alongside other parts where yeah, it can’t count.

All caught up in that is whether it’s going to rewrite employment structures or funding structures. You also raised the murky question of, what is mathematical knowledge? And the roles that these workers will be doing as well. It’s all of that wrapped into one.

I think that tracks broadly with the rise of AI in every field. Where you can just add horsepower or compute to a problem, and there’s some kind of verifiability, it seems like the models continue to get better. Everything in the middle where you might need some world knowledge or the models might need some actual intelligence about the world itself, they seem to struggle.

Those parts of math, at least reported out in your piece and what the labs are talking about, seem to be almost entirely self-contained theoretical problems. That’s where the models can generate a proof or solve a problem that no one’s been able to solve, and then try to verify that that has existed and they can just run it again and again and again.

That brings us, I think, to May of this year where an internal OpenAI model, which we have not really seen, disproved the unit distance conjecture, which is an 80-year-old problem. And then just recently we heard about Astra from OpenAI. Astra is the one where it seemed like the switch flipped, and everyone decided it was an existential crisis.

What did Astra achieve, and why is it a big deal?

I’m also pretty sure that Astra was probably behind the early one as well. OpenAI just listed it as an unnamed internal model. It’s probably Astra. They didn’t answer me when I asked.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近