Top
Best
New

Posted by auggierose 2 days ago

Why do we need human mathematicians anymore?(terrytao.wordpress.com)
291 points | 374 commentspage 4
ed_elliott_asc 2 days ago|
I’m just shocked by how “intelligent” people believe a token guessing system can be
dist-epoch 2 days ago||
Pick one:

- solving "frontier" math problems requires intelligence (by human or AI)

- solving "frontier" math problems is dumb statistical prediction of next token (by human or AI)

auggierose 2 days ago|||
I am shocked how people can deny that solving Navier Stokes requires some sort of intelligence. Even Doctorow talks of "brute-forcing" a solution. Brute-forcing leads to combinatorial explosion, so there must be something more going on here. Otherwise you could just put this problem into an automated theorem prover (we've had those forever, they are actually just brute-forcing it).
robotpepi 2 days ago||
> Otherwise you could just put this problem into an automated theorem prover (we've had those forever, they are actually just brute-forcing it).

There are many automatic theorem provers that do very clever stuff, just as the underlying theroy describes.

> I am shocked how people can deny that solving Navier Stokes requires some sort of intelligence.

It is absurd to waste time discussing whether it is inteligent or not. It is just an algorithm, we know how it works, and it does exactly what we expect it to do. LLMs are not magical things. The main difference is the scale: for Navier-Stokes they spent in 3 days more money that the whole mathematical community over the last 20 years easily.

By the way, I'm not saying that LLM's are useless, that I'm anti-AI or anything like that.

auggierose 2 days ago|||
Intelligence does not seem to be magic either, as LLMs are proving now. It is indeed a waste of time to argue that LLMs are not intelligent in their own way, they obviously are. If Navier Stokes doesn't convince you, nothing will.

I just used a £89 Codex subscription to do very intelligent things with it, stuff that I would have had to sit down and ponder and work on for quite a while, and I have a PhD in that. I didn't need to do anything special except explaining the problem(s) to the AI, and my theory of it so far. It took it from there. If that is not intelligence, nothing is.

gjm11 1 day ago|||
> It is just an algorithm, we know how it works,

We know what calculations it does. We have some hazy idea of some bits of how those calculations lead to something that at least somewhat resembles intelligent behaviour. But that's a far cry from actually knowing how it works.

For instance, suppose you give one of today's frontier models some of those chain-of-cubes rotation puzzles (the sort that infamously men are about 1sd better at than women, statistically speaking). How well will it do? I have absolutely no idea and I'm quite sure that a more detailed understanding of the transformer architecture would not make my guesses any better. (Actually, I do kinda have some guesses but they're based on a vague notion about how the models might be partitioned between vision-y bits and language-y bits, and it's very possible that that notion is out of date.)

> it does exactly what we expect it to do

Were you, let's say 6 months ago, expecting it to resolve one of the Millennium Prize problems?

(I do agree that it is more productive to ask "what can and can't they do?" than "should we classify that as intelligent or not?".)

> for Navier-Stokes they spent in 3 days more money than the whole mathematical community over the last 20 years easily.

Are you sure?

(The numbers I've heard, which I admittedly have no very strong reason to trust, don't seem that way to me.)

robotpepi 1 day ago||
> Were you, let's say 6 months ago, expecting it to resolve one of the Millennium Prize problems?

I didn't expect them to throw millions of dollars at each famous math problem. But one year ago we already had LLMs that solved IMO problems, no?

> Are you sure? (The numbers I've heard, which I admittedly have no very strong reason to trust, don't seem that way to me.)

Math has very little founding compared to other science domains. Also, if you filter mathematicians by specialization in PDE and that have worked on Navier-Stokes, then you end up with a very niche community.

> For instance, suppose you give one of today's frontier models some of those chain-of-cubes rotation puzzles. How well will it do?

I feel like this is not the correct way of thinking about it. We can also ask, for instance, how well a state-of-the-art algorithm for the salesman problem works on a particular graph topology. People do PhD thesis on topics like that, so the answer is not obvious at all. For LLMs we still don't have a curated theory that explains what they're good/bad at, and that you don't see how to extract an answer from the definitions is no surprise since this is obviously not an easy problem. But all this is normal because this is a rather new topic (models of this scale appeared when? 3 years ago? That's nothing for science).

Anthropomorphizing LLMs has added so much noise to this discussion.

gjm11 1 day ago||
Yes, one year ago we had LLM-based AI systems solving some IMO problems. My impression is that most observers at that time didn't expect them to be solving Millennium Prize problems within a year.

> Math has very little funding compared to other science domains.

True. But to whatever extent the numbers I've seen are correct, for the whole mathematical community to have spent less on Navier-Stokes than OpenAI did -- even if we value the tokens they spent at something like market rate rather than at what the compute actually costs them (which might be right since any capacity they use internally can't be sold to customers) -- the average number of mathematicians working on Navier-Stokes since 2000 would need to be somewhere around four (depending of course on how well paid they are), and that seems too low to me.

> I feel like this is not the correct way of thinking about it.

It seems to me that if you say "It is absurd to waste time discussing whether it is intelligent or not. It is just an algorithm, we know how it works, and it does exactly what we expect it to do." then this only makes any sense if your "knowing how it works" and "what we expect it to do" enable you to predict what it can and can't do.

(I repeat that I agree that what matters is what it can do, not whether we choose to apply the term "intelligent" to it. But unless I misunderstood you were saying somewhat more than that.)

> Anthropomorphizing LLMs has added so much noise to this discussion.

I think sometimes it helps, sometimes it hurts, and sometimes it's indifferent, because LLMs are like us in some ways and unlike us in some ways. (The same goes for many other things, but LLMs are much more like us in some important ways than any other human-made artefacts.)

hardbass 1 day ago|||
I am shocked by how intelligent people seem.to believe in some form of dualism or supernatural element to thought and intelligence.
ComplexSystems 2 days ago||
Turns out what you call "intelligence" was never needed to do math.
goatlover 2 days ago||
It was needed before the LLMs and training data existed in digital form. Euclid, Gauss, Turing didn't have that benefit.
BigTTYGothGF 1 day ago||
Lots of deans and college presidents asking themselves the same question.
YeGoblynQueenne 1 day ago||
I get the feeling that mathematicians are needlessly panicking because they don't really understand how AI works. They see the results, but they haven't thought enough about the methodology and so they don't have a clear picture of the true capabilities of thsoe systems.

For the n'th time: the recent successes of AI in mathematics are the result of a brute-force attack. See the proof for Navier-Stokes: 10k agents running for 88 hours; that's ~100 GPU years. How many human-years were invested in solving the same problem, before they were overtaken in the last few days by an AI? 90? Not even: that's just the time since Jeal Leray's statement of the problem in 1934. 26, if you want to count the time since 2000 when the Clay Institute named it as one of its Millennium Prize problems. But how much time have human brains spent working on the problem in either of those time periods? How many mathematicians have worked on the problem? 10k? Not likely.

And all that's without even considering whether the AI based its proof on carelessly shared work by the humans. Or rather, yes, let's consider that: it totally did.

Further. There have been several results in mathematics produced by AI but we have no information on how many attempts were made to produce similar results that failed. Because we don't have this information we cannot estimate the true capabilities of AI.

Yet we can observe that, for example, out of the six Millennium Prize Problems remaining open before the claim of a solution of Navier-Stokes existence and smoothness, only one (the aforementioned) was solved by an AI. We can assume that the AI companies (more than one) tried and failed to solve the others. We can even guess that they previously tried, and failed, to solve Navier Stokes itself, and only succeeded once the progress made by Buckmaster and Alpöge was in the training data [1]. That's a success rate of one out of six, or ~17%. That's what's gonna solve all of maths and destroy the tradition of mathematics? A success rate of 17%? Well, grab a Snickers 'cause we're gonna be waiting for some time!

Moreover. If we include in the list the Poincaré conjecture, proved by Grigori Perelman, who is a human, that's a score of AI 1-1 Humans. And that's being gracious: we have one Millennium Problem fully solved by humans, one solved partly by humans with a last-mile solution by AI. We have thousands of problems solved by humans in the last 2k years and how many by AI? A couple dozen? Oooh scary!

- Hey Hal! Prove that P ≠ NP!

- I'm sorry Dave. I can't do that.

What I'm trying to say, without the snark (sorry): Panic if you will, but the machines are not yet taking over. If you're panicking, panic for what you believe they will be able to do in the future. Because they certainly can't do hat in the present. They can't solve "all of mathematics" (whatever that means).

______________

[1] Yes it was. Buckmaster reported that he turned off the option to train on his data in July, after working on the problem with Alpöge for a year since September 2025. OpenAI claimed a solution in September, a month after they had stopped hoovering up Buckmaster's data. They had plenty of time to train on his data. Ask for references if you want them because I don't have them handy right now.

lelanthran 1 day ago||
> I get the feeling that mathematicians are needlessly panicking because they don't really understand how AI works.

I sunno if mathematicians would be having problems understanding how matrix multiplication, backprop, sigmoid functions, attention, embedding distances, probabilities, etc work.

As a group, they are probably more likely to understand it than everyone else.

YeGoblynQueenne 1 day ago||
That's a bit like saying that a physicist is more like to understand how a car works than anyone else because they understand all the principles of an internal combustion engine. And yet, curiously, when we take our car to the garage the person fixing it does not tend to have a physics degree.

Wanna guess why? I'm too tired now to expand the argument properly but basically understanding the components of a complex system doesn't mean you understand the principles of the system. A mathematician who is not an expert in AI has no reason to be particularly capable of understanding how AI works, i.e. how all the maths that go into creating an AI system come together to create. An AI system.

gjm11 1 day ago||
> How many human-years were invested in solving the same problem, before they were overtaken in the last few days by an AI? 90? Not even: that's just the time since Jeal Leray's statement of the problem in 1934. 26, if you want to count the time since 2000 when the Clay Institute named it as one of its Millennium Prize problems.

I hope this isn't actually news to you, but: There is more than one human. There is even more than one mathematician.

If there happen to have been as many as four humans working on Navier-Stokes at any given time since the year 2000, then that's more human-years applied to the problem than agent-years.

> How many mathematicians have worked on the problem? 10k? Not likely.

You don't get to count the factor of 10k once when working out how many agent-years OpenAI gave to the problem and again when demanding that for parity there would need to have been 10k mathematicians on it.

> And all that's without even considering whether the AI based its proof on carelessly shared work by the humans. Or rather, yes, let's consider that: it totally did.

Let's suppose that indeed what Buckmaster and Alpöge had done was in the model's training data. Well, it didn't enable Buckmaster and Alpöge to solve the problem for Navier-Stokes (they could only do Euler), and it did enable OpenAI's model to do that.

Also: we don't actually know that what they'd done was in the training data; the latest bits of what they'd done that could plausibly have been in the training data were from before when Buckmaster said they progressed from preliminaries ("We worked through the literature and upgraded various preliminary results") to actually making substantial progress on the problem ("This was until about a month ago, when we had real progress"); and from what Buckmaster wrote it sure seems like a lot of the Buckmaster/Alpöge progress was in fact done by LLMs. (E.g., Buckmaster says that he and Alpöge have been working frantically to try to understand the proof for their Euler solution. That sounds to me much more like "an LLM did this thing" than "we figured out all the hard bits and the LLM did nothing more than filling in a few details".)

Buckmaster's own account of things is that all the really clever ideas were those of Córdoba and Martínez-Zoroa. (Which are already out there in the open literature, and there is nothing remotely improper about making use of them.) And my understanding (but, note, I am not an expert on fluid dynamics or PDEs and I could be wrong) is that actually the OpenAI model's construction is quite different from that of C&MZ. On what basis are you confident that "the AI based its proof on" what B&A did?

(For the avoidance of doubt: I am not arguing that what OpenAI did was OK. Even if they actually didn't train at all on any of the Buckmaster/Alpöge chats, it's very much not good professional ethics to hear that someone else is working on something and rush to try to scoop them, and there is absolutely no question that they did that. The question here is how impressed we should be by the model's mathematical prowess.)

> A success rate of 17%?

A success rate of 17% on problems of this difficulty and significance is something that for any human being would be a career-defining triumph.

> We have thousands of problems solved by humans in the last 2k years and how many by AI? A couple dozen? Oooh scary!

That would be a more convincing argument if the AIs, like the humans, had been around and trying to solve those problems for the last 2k years. However, as you might have noticed, the state of the art in AI was rather primitive 2000 years ago.

YeGoblynQueenne 1 day ago|||
>> If there happen to have been as many as four humans working on Navier-Stokes at any given time since the year 2000, then that's more human-years applied to the problem than agent-years.

My bad for not showing my work and inadvertently leading you down the garden path, but the "~100 agent-years" calculation goes like this:

10,000 agents * 88 hours = 880,000 agent-hours

88,000 agent-hours / 24 hours = 36,666.7 agent-days

36,666.7 agent-days / 365 days = 100.5 agent-years.

That's what you get for working 24 hours a day, 7 days a week, 365 days a year. Realistically speaking, that's not a work schedule any human can follow.

It's hard to make a realistic estimate because normally even a very dedicated mathematician will not be working exclusively on one problem all their waking time, or even all their working time. But, let's ignore this and assume a pretty standard work schedule of 8 working hours, five working days a week, and 52 working weeks a year.

Now, that's:

8 hours * 5 days = 40 working hours a week

40 hours * 52 weeks a year = 2080 hours a year

880,000 agent-hours / 4 humans = 220,000 hours per human

220,000 hours per human / 2080 hours a year = ~105.8 years

To clarify, that's how I estimate the number of years it would take a mathematician to do a quarter of the work of the 10k OpenAI agents if that mathematician worked only on solving Navier-Stokes and did nothing else in their entire career.

That's just not a realistic work schedule for any human. You can adjust the working hours if you want but I don't believe you'll get any realistic estimate. Don't forget that most academics' careers last around 30 years from PhD to Professor Emeritus. If you want a more realistic estimate of how much time it would take how many humans to do the work of the 10k OpenAI agents, you can start from that assumption and work your way up from that.

>> That would be a more convincing argument if the AIs, like the humans, had been around and trying to solve those problems for the last 2k years. However, as you might have noticed, the state of the art in AI was rather primitive 2000 years ago.

Sure. But the thing is agents can run 24/7, 365/365 in parallel and as you see above they can cover 2000 years of human work in much less time. I'm not going to estimate how much because the only bottleneck is the amount of compute and money that an AI company wishes to spend, and that depends on their motivation to solve a particular problem. However, with sufficient motivation 2k years of human research (keeping mind that's not 2k years of continuous work) can be covered in a few ... months? Probably.

gjm11 1 day ago||
> My bad for not showing my work and inadvertently leading you down the garden path

The problem isn't that you didn't show your work, it's that your work was wrong.

I entirely agree with your calculation that 10k agents for 88 hours is about 100 agent-years if we assume 24/7/365 operation. That's not what I was disagreeing with.

But then you said "How many human-years ...?" followed by estimating not the number of human-years that have gone into the problem but merely the number of years.

You can compare elapsed years for humans (26) and elapsed years for AI systems (about 0.01). You can compare agent-years (about 100) and human-years (26 times the average number of humans working on Navier-Stokes at any given time). Either of those is defensible.

But it makes absolutely no sense at all to compare agent-years for the AIs and elapsed years for the humans. Which is what you did.

If a typical human mathematician works 2000 hours a year (actual human mathematicians generally find that they can't do 8 hours a day of focused hard intellectual work, but I think we should count some of their "percolation time" too) then that's about 6 human-years per mathematician. So to get the same amount of mathematician-work as agent-work the average number of mathematicians you need to have been on the job is about 100/6, or about 16.

So when you wrote

> How many mathematicians have worked on the problem? 10k? Not likely.

the 10k figure was a total irrelevance. The number it would actually have to have been is about 16.

(My earlier "as many as four" ignored the fact that humans don't work 24/7/365, as you point out. But my point is that however you slice it the relevant number is more like four than it is like 10,000.)

My guess, for what it's worth is that that is roughly the order of magnitude of the number of human mathematicians working primarily on things that could be classified as "trying to make progress toward resolving the Navier-Stokes problem" during that time. I wouldn't be surprised if the actual figure were 3x bigger or 3x smaller. It probably depends on how broadly you interpret "trying to make progress toward resolving the Navier-Stokes problem", and one important difference is that all those human mathematicians leave behind them a trail of papers proving things that, whether or not they end up on the path to Navier-Stokes, may turn out to be useful later, whereas if OpenAI's agent swarm proved a lot of useful theorems along the way most of them never got published.

I don't, of course, disagree that it's possible for an AI company to put a lot of AI agents to work on a problem, but I'm not sure how that makes what they can do less impressive. The fact that you can do that has always been a major part of why AI could be such a big deal. "A country of geniuses in a datacentre" is the kind of thing people have said; we aren't quite there yet, but the "country" part is as important as the "geniuses" part.

YeGoblynQueenne 21 hours ago||
I'm sorry but I'm not sure I understand your argument. I think you're saying I'm comparing apples to oranges. I'm not: I'm comparing apples to apples and oranges to oranges. These are two different questions:

>> But how much time have human brains spent working on the problem in either of those time periods? How many mathematicians have worked on the problem? 10k?

So neither 10k humans worked on Navier-Stokes, nor has any human spent a century of non-stop work on it.

But I could have made the point more clear maybe.

>> I don't, of course, disagree that it's possible for an AI company to put a lot of AI agents to work on a problem, but I'm not sure how that makes what they can do less impressive. The fact that you can do that has always been a major part of why AI could be such a big deal. "A country of geniuses in a datacentre" is the kind of thing people have said; we aren't quite there yet, but the "country" part is as important as the "geniuses" part.

Yes, I see your point, but those are not geniuses. Grigori Perelman proved the Poincaré conjecture alone, though as he has emphasised his work was based on advances made by others, particularly Richard S. Hamilton. That we can call a genius: a single man who solves one of the most interesting problems in all of mathematics building on the work of his predecessors. 10k agents that search blindly and find a result by luck (or by stealing it), I don't agree we can call "genius". That's what I call "brute force". Anyone who wants to call OpenAI's agents "a country of geniuses" has first to deal with the fact that they look a lot like monkeys on typewriters.

FrustratedMonky 2 days ago||
Because the AI needs new stuff to train on?
jdw64 2 days ago||
I sometimes think about this: to grasp the fundamentals of knowledge and complex phenomena, perhaps we need an external system rather than human knowledge systems. By that logic, maybe we need AI, which can handle far greater complexity.

Human capability, when you think about it, is complex. Why is Newton praised as being so damn great? He established the law of universal gravitation, F=ma. Why is that such a big deal?

He distilled countless phenomena in an open system into a single mathematical formula.

What makes it great is that he found common state variables and relationships across entirely different phenomena like falling objects, planetary motion, collisions, and artillery trajectories.

But does F=ma hold true for the entire macroscopic world? No. There are various conditions and specific situations in motion, but within most scenarios and a certain range of approximation, it outputs values that are useful to humans.

Why is the Schrödinger equation so great? Because it turned the time evolution of quantum states into a calculable mathematical law.

Human thought is essentially creating a closed system by deciding what to cut out and what to keep from the infinite degrees of freedom in reality. Academia is what reinforces that closed system.

A great theory is great not because it perfectly replicates reality, but because it compresses the immense complexity of reality into a small, closed formal system while still managing to explain a multitude of phenomena.

In that process, it feels like human thought and progress are shifting into a different framework. What LLMs do well is primarily exploring within the ontology and representation space that humans have already built.

I think there are two broad categories of discovery: One is forming a new closed system, and the other is connecting fragmented knowledge within that closed system. I feel that the vast majority of research focuses on the latter.

What LLMs excel at is finding unvisited points within a given representation space. This is typically the process through which master's and PhD students connect dots, build their skills, and form their own mental models. But the logic behind criticizing LLMs seems to be that they eliminate the very work these graduate students need to do in order to grow.

However, looking at it from another angle, perhaps our current knowledge systems and classifications have reached a limit, suggesting that we might actually need a completely new classification and knowledge system.

What is the core principle of an LLM? It's predicting the probability of the next sequence.

Let's say you type the word "cat". Cat - is cute (90%), want to eat it (6%), furry (4%). Because "is cute" has the highest probability, the next sequence proceeds in that direction.

Within this framework, human knowledge and logic largely operate the same way. Once an initial logical proposition is established, we follow it up with whatever makes logical sense next. From that perspective, I think LLMs will actually do this better.

But what is it that LLMs cannot do right now? They cannot create that initial logical proposition. I believe they lack the ability to carve out a closed system from an open system.

Stacking logic step-by-step within a closed system—LLMs do this exceptionally well. But whether that constitutes true "intelligence" is a different matter.

I feel that being logical does not necessarily equate to having intelligence.

Humans preserve and create different mental models and knowledge systems within an open system. Just as your thoughts differ from mine, LLMs lack the ability to form these distinct mental models.

If so, within these limits, what humans must ultimately do is construct the logical frameworks that LLMs can then fill in. Perhaps a new kind of logic dedicated to designing these frameworks will become the next major trend.

Viewed from this perspective, I have no idea if we are in a mere technological transition or something else entirely. Or whether it is even correct to say humans are strictly necessary to build that framework. Maybe my learning is just lacking.

ck2 2 days ago||
waiting for the xkcd for that

I always think of this one but I bet there's better

* https://m.xkcd.com/435/

card_zero 2 days ago|
Got to be philosophical about it all.
vixen99 1 day ago||
'There are zero examples of any intelligent species which is vastly more capable than another species, yet surrenders control ... to the less capable species'.

True with genuine species. But we should take note of the plentiful counter examples within human societies. How about politicians and our method of choosing those people to whom we delegate the most critical decisions regarding our and the Earth's future? We select politicians mostly either by rote or via their persuasive rhetoric, their general personality & likeability and probably least of all by their intellectual capability or indeed general capability in too many cases. That is not to say intellectuals are necessarily any better at the job. There are numerous other examples in human organizations as we know, sometimes to our cost. Truth is that we cannot even agree on how, as a species together with the other life forms on Earth, we can all 'flourish' though there are lots of great examples working in local environments.

dbg31415 2 days ago||
Funny... for me, this article was right next to the link to this current thread.

AI chatbots give wrong answers to financial queries 'most of the time' (ft.com) // https://www.ft.com/content/c0cd359d-df84-4208-a789-ffa864b43...

johnea 1 day ago||
We need human mathematicians so that someone is creating the data that the models are trained with...
podocarp 1 day ago|
> this is a loss of control from incumbents in a scientific field

Lol, the ultimate delusion, laughing at people in the flood zone and not seeing the tsunami... What goes around comes around

More comments...