Welcome to the AI crisis in math

AI Summary
В интервью с Робертом Хартом, AI-репортером The Verge, обсуждается влияние искусственного интеллекта на математику и кризис, с которым сталкиваются ведущие математики. OpenAI представила решения для давних математических проблем, что вызвало бурное обсуждение в сообществе и поставило под сомнение необходимость обучения новых поколений математиков, если AI способен решать сложные задачи. Это поднимает важные вопросы о будущем математики как академической дисциплины и о возможных маркетинговых мотивах AI-лабораторий.
Today on Decoder, I’m talking with Robert Hart, The Verge’s London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis many lead mathematicians are having about it.
OpenAI just published a set of solutions to longstanding problems in math that went off like a bombshell in the field. It caused a huge debate in the math community, and Rob spent some time talking to some of the most accomplished mathematicians of our time about it.
It’s funny that AI systems are all still pretty bad at elementary school arithmetic, but getting increasingly good at very high-end abstract math. That raises some big questions for the field of advanced math.
If AI can do math of this caliber, does that mean AI labs can transfer those skills to other domains? What good are academic grants and university programs training new generations of human mathematicians to identify new problems as they try to solve existing ones, if frontier models simply answer all the outstanding questions?
What if all this attention around math is just a big marketing exercise for frontier AI labs, which couldn’t care less what happens to one of the oldest and most fundamental academic disciplines there is?
There’s a lot here, and Robert has talked to a lot of people with a lot of views on all of it.
Okay: Verge AI reporter Robert Hart on what AI is doing to math. Here we go.
This interview has been lightly edited for length and clarity.
Robert Hart, you’re our London-based AI reporter here at The Verge. Welcome to Decoder.
Thank you for having me.
I am very excited to talk to you. There’s a lot going on in particular with AI and math that you recently dove into. You spoke to a lot of leading mathematicians about the crisis in mathematics due to AI.
It feels like a lot, and also like there’s a lot yet to know and discover about the interaction of these two things. A full existential crisis, which is pure Decoder bait. Broadly tell us what’s going on.
I think “a lot” sums it up quite well. Basically a bit of an existential crisis within, “what is mathematics? What are mathematicians doing, and what is the role of mathematicians going forward?”
A lot of that has been spurred by a phrase transition in what AI is capable of that has exploded in the last six months to a year. AI went from being very terrible to seemingly genuinely quite good at a professional level in a very short space of time. It’s a lot of what these other fields have been struggling to deal with for the last five years in a very compressed period of time.
I would put that next to software engineering. We’ve been living through the AI crisis in software engineering for some amount of time. But as recently as 2024, even last year, the conventional wisdom was that AI models were particularly bad at math. The famous example is that these models could not count the number of R’s in the word strawberry. Even just counting eluded them.
What has happened to make them better at math? Are they still bad at general arithmetic and they’re good at advanced math, or is it something in between?
They are still truly, truly terrible at some areas of math. I did check, they can do strawberries now. I think someone’s tweaked it.
I think strawberry’s hard-coded. I want to be very clear, my conspiracy theory is that the strawberry thing is hard-coded into all the models.
I think so too. That is a conspiracy I’ll buy into.
But yeah, it’s still terrible at those kinds of things — math, arithmetic, even the days of the week. My boyfriend was saying the other day, “It keeps thinking it’s Wednesday. It’s not Wednesday.” Or time. Elissa Welle for us a few months ago wrote that ChatGPT can’t tell time. Still can’t. That’s not all of math.
So there’s this disconnect. To be good at math, you’ve got to be good at counting, or adding, or multiplying. A lot of it is actually reasoning. If you look at academic math papers, a lot of the time you won’t see numbers, which sums that one up, I think. So they’re still terrible, but they’re now also very good at this other part.
As to why, at some point you reach a critical mass of what these systems can do. We saw it with writing, we’ve seen it with programming. They’re very good at forging connections between different areas, applying old methods in new ways, those kinds of things. It appears that the newer models they’re training have apparently reached that level where it clicks, and now it can do math.
It’s important to say as well that we speak of math as a unitary discipline, especially from the outside. But imagine, say, biology. You’ve got something that would range from literally watching animals and describing behavior all the way through to cellular mechanisms and bioche
