Top
Best
New

Posted by 6bitquant 1 hour ago

The Mathocalypse(scottaaronson.blog)
118 points | 108 comments
zaxioms 1 minute ago|
I'm a PhD student in CS. While I think these results are rather cool, it makes me terrified that the skills developed by the PhD will ultimately be worthless. I'm not quite sure what to do. Any thoughts from people in similar positions?
ks2048 1 hour ago||
> It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results

> Basically the paper is so horribly written that it’s impossible to read it without AI help

That's interesting and haven't seen this in all the coverage of this event.

It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.

TheOtherHobbes 29 minutes ago||
Math proofs need to produce the correct output correctly, which is not quite the same thing.

This looks like an AI IPO PR powerplay, because at this point the proofs haven't been checked and it may not be possible for a human to check them - because proofs should be clear, not horribly written and noisy.

The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.

You want the path through the maze to be as short as possible and the map to be as clear as possible.

This sounds like the opposite. There may be a genuine path through the maze, but if it's too convoluted and takes too long it will be impossible to confirm.

I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.

I suspect that's possible without tripping over the halting problem. (But I can't prove it.)

FloorEgg 1 minute ago|||
If intelligence is compression, and these models are a different form of lesser intelligence than human, but being scaled up to brute force problems, then it makes sense the artifacts that produce (the proofs) would have worse compression than a human proof would.

In other domains I have seen first hand overwhelming evidence of how things that cause the AI to make mistakes also cause humans to make the same mistakes.

I wonder if the proofs being produced that are hard for humans to interpret are also hard for other LLMs to interpret.

In other words, I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.

It kind of an explicit example of how the LLMs can be materially less intelligent than people, but still be more productive through scaling, and yet they also can't replace people because they are a categorically different kind of intelligence. It's like all the AI debates compressed into one example showing countwr-intuitive answers.

eadler 9 minutes ago||||
That reminds me of this paper:

Chow, T. Y. (2008). A beginner’s guide to forcing (arXiv:0712.1320). arXiv. https://doi.org/10.48550/arXiv.0712.1320

> “All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves. The proofs should be “natural” in Donald Newman’s sense [13]:

> This term . . . is introduced to mean not having any ad hoc constructions or brilliancies. A “natural” proof, then, is one which proves itself, one available to the “common mathematician in the streets.””

sebzim4500 11 minutes ago||||
Surely by the time of the IPO we will know whether the main results are correct, if only because a different AI will have produced a lean proof or found a logical flaw (the second case would be hard to verify but probably not impossible).

Also from what I can tell from the few fields I understand, the proofs aren't that long or complicated they are just terribly written.

Octoth0rpe 10 minutes ago||||
> A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.

https://en.wikipedia.org/wiki/Inter-universal_Teichmüller_th... seems like a counterpoint, but IANAM. (I am likely cherrypicking the far end of the bell curve re: straightforward here)

pizza234 11 minutes ago||||
The post says there's a Lean certificate for this and other proofs ("some [...] not all of them").

> This looks like an AI IPO PR powerplay,

Interestingly, the post has actually also an argument for this:

> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.

> So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.”

smcg 2 minutes ago||
It's on OpenAI and Anthropic to prove that they obtained these results legitimately and credited all researchers who deserve credit. They do not get the benefit of the doubt.
caaqil 10 minutes ago|||
We should consider the possibility that at some abstraction levels, we can safely stop chasing "clarity" or "coherence" which is circularly defined in such a way that it's capped by human processing power.

Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. Mathematicians will get there.

ssfdg 16 minutes ago|||
This proof dump reminds me of the glut of low-quality drive-by PRs overwhelming open-source repos.
dormento 9 minutes ago||
Its like infinite summer of code, but for math. Must be annoying.
piker 55 minutes ago|||
It also aligns with the fear that these proofs present a risk to the ecosystem by out-competing attempts at more human-readable proofs. Perhaps though we end up with more math influencers who edit and annotate these proofs to bring them back to us.
whatshisface 22 minutes ago|||
The ecosystem is (ahem) gated by hiring committees. There is no risk of AI replacement from the inside. "Replacement" is not even a possible movement. The funding for mathematics worldwide comes mostly from endowments, which are investment pools.
bobajeff 36 minutes ago||||
I think that's ultimately a good thing. As proofs weren't supposed to be the point as stated by William Thurston long ago. Maybe now the focus can be more on better explanations and creating tools for growing understanding and intuition.
cowlevel 10 minutes ago||
Good explanations should take the form of human-understandable proofs.
rrr_oh_man 35 minutes ago|||
Vibe mathing
jltsiren 16 minutes ago|||
Isn't that just the default experience with AI these days? In small enough scale, AI models can express their ideas clearly. But the larger and more complex the ideas are, the less suitable the outputs are for human consumption. I guess AI models think too different from humans, and nobody has trained them to communicate complex ideas in the way human experts in that particular topic expect.
spelunker 7 minutes ago|||
I see many parallels to genAI-assisted code development. Not surprising I think.
aaroninsf 41 minutes ago||
Serious question:

Why would anyone believe this (also) is not simply example N+1 of this is the worst it will ever be, as opposed to recognizing this as what will almost certainly prove to be an awkward moment, soon to be replaced by another order of magnitude of cleaner, clearer, more intelligible, etc.?

Ximm's Law: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon.

Kotlopou 1 minute ago|||
In that case, one would expect to see some progress in this direction, but AFAICT that hasn't shown up yet? If anything, it's getting worse, though that could just be the increasing scale and decreasing cleanup efforts.

Already the unit distance proof was substantially human-edited (per Thomas Bloom). Then with the ten problems from Astra you started getting the citation issues. Then Navier-Stokes was a rushed 160 pages with barely any citations, and some of the related papers were called (by their "authors") the ugliest mess they've ever seen.

And now here we are. At least it seems that mathematical ability and communication with a mathematical audience are independent skills, and progress in the first does not imply the second.

This doesn't surprise me much, given two analogies: 1) many smart people are nonetheless horrible lecturers. (You can't quite get the opposite extreme, since to explain math well you have to be able to do it.) 2) AI writing in general hasn't improved. The models have annoying verbal tics ("honestly") and have no sense of which part of what they say is obvious and which is relevant.

devin 32 minutes ago|||
Devin's Law: every defense of AI which rests on "it will get better, trust me" is in many ways indistinguishable from 2010s crypto hype or "level 5 self driving is right around the corner"
usrnm 5 minutes ago||
1) Predicting the future is hard, but so far everyone who was saying that it would get better turned out to be right. It is getting better 2) Waymo exists
cgio 1 minute ago||
I thought that was from the outset the intent of the Hilbert program, to automate mathematics. And mathematicians were behind it. Cannot see why they would be concerned when a different way to do the same, not subject to Gödel incompleteness, is working out. Maybe the frustration is that they were not the ones building it.
nostrademons 1 hour ago||
As a side note, you can tell this wasn't written by an AI by the first sentence:

> mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!

My 8yo talks exactly like that. I could totally imagine him saying this, the same way, at the dining room table.

I asked ChatGPT "pretend you're an 8/9 year old today. how would you insult your mom about having her job be replaced by an AI?", and the responses it offered were:

> “Mom, AI took your job because apparently even robots were like, ‘Yeah… we can do this better.’”

> “Mom, congratulations! You got replaced by a computer. Even Siri has a job now and you don’t!”

> “Mom, AI took your job? Dang. I guess even a robot looked at your work and said, ‘I got this.’”

> “Don’t worry, Mom. You can still be useful… like teaching the AI how to make my lunch.”

All of these seem to have a vaguely Millennial flavor, aside from being pretty awkward and mechanical roasts. Trust the children and linguistic drift to be the best AI detector.

posnet 51 minutes ago||
[dead]
john_strinlai 1 hour ago||
[flagged]
softwaredoug 37 minutes ago||
Aren’t there dozens of proofs of the Pythagorean theorem? The goal isn’t to just “prove” but create something well written and intuitive to the average practitioner. And by gaining a deeper understanding we can ask better questions.
GuB-42 11 minutes ago||
Something that often comes out is "it is about the journey, not the destination".

Many math problems are practically useless if you only care about the answer, the millennium prize about the Navier-Stokes equation is such a problem. The solution makes no physical sense, real life fluids don't follow the Navier-Stokes equations in such extreme conditions. But in the process of finding the solution, we may get insight into what will end up being really useful. The big mess that OpenAI produced is the solution no one really cared about, but it didn't deliver much of what people actually wanted.

One reason it is sometimes seen negatively despite being at least something is that it broke the incentive. Without the million dollar prize and with only the privilege of being second, people are much less likely to go for the insightful solution.

soVeryTired 19 minutes ago|||
But up until now, the mathematics community has valued the "prove" part much more highly than the "deliver an insight" part. Mostly because with a little work they went hand in hand.

And going from zero proofs to one proof (even a sloppy one) is a big deal regardless of whether it was written by AI or a human.

softwaredoug 6 minutes ago||
To be frank, the obsession with being first, and not making research accessible, has always held academia back
WD-42 26 minutes ago||
No, haven’t you heard? Since the AI bubble began we’ve collectively decided that outcomes are all that matter. /s
p0w3n3d 24 minutes ago||
Recently I asked ai to tell my daughter how to quickly calculate 11^2 12^2 etc but the outcome it gave was horrendous. I quickly shut it down and gave her better ideas
phoghed 20 minutes ago||
Hi, I’d like to signal that I’m part of your in-group. One time I used AI and it sucked. Every time someone says it’s good, they are lying because they are shills. Upvotes to the left please.
smcg 5 minutes ago||
How do we know that these "internal models" are not just half computer and half a giant team of mathematicians? How do we know that OpenAI actually came up with these solutions and didn't steal them from outside researchers?
thejokeisonme 2 minutes ago||
How would these ideas be available to steal?
UltraSane 2 minutes ago|||
It would be extremely unlikely human mathematicians able to solve these kinds of problems would accept not getting credit that would set them for life professionally.

Also lean proofs are notoriously tedious and slow to write so this level of output is very likely to be from LLMs.

runarberg 4 minutes ago||
Until this is replicated, we don’t.
an0malous 1 hour ago||
> But it also appears that no human has understood just about any of these proofs yet

Has anyone verified any of the proofs produced by OpenAI or is everyone just assuming that it just be true because the Lean code checks out? Couldn’t the Lean code just be formulated incorrectly?

prof-dr-ir 21 minutes ago||
It's a mixed bag I think.

For example, the statement of e.g. Fermat's last theorem in Lean should be understandable to anyone who played The Natural Number Game [0] and knows a bit of mathematics and programming. For the proof, you trust the compiler.

The statement of other theorems can be much more delicate, and the Lean formalization may require an extensive introductory section which will need to be carefully checked.

Then there are the cases where no Lean formalization is currently available, and all we have right now is an often impenetrable pdf in the OpenAI repo. I would not at all be surprised if some of those claims contained logical gaps.

Time will surely tell, but there are certainly doubts and lots people are very busy checking these results.

[0] https://adam.math.hhu.de/#/g/leanprover-community/nng4

nperez19 1 hour ago||
There's an entire paper claiming that many of these AI-generated Lean proofs are formulated incorrectly / mistranslated: https://arxiv.org/abs/2610.08144
nsingh2 21 minutes ago|||
Note that paper is saying that the lean proof and the natural language proof do not necessarily coincide. It is not saying that the lean proof is wrong, just that the lean proof does not necessarily mean the natural language proof is correct.
macleginn 13 minutes ago||
The thing is, you often see people saying, ‘They have a Lean cert, so it has to be correct, even if I don't understand it.’
sebzim4500 54 seconds ago||
They are right? The lean proof is correct. It's the natural language proof that potentially isn't (or at least it isn't identically structured to the lean proof)
sigmar 14 minutes ago||||
that paper isn't saying that. why are there so many single digit karma accounts misrepresenting that paper?
ajjenkins 33 minutes ago||
The line about “understanding the aliens” reminds me of Ted Chiang’s short story The Evolution of Human Science (2000).

Highly recommend reading it. Very prescient for something written 26 years ago.

https://gwern.net/doc/fiction/science-fiction/2000-chiang.pd...

nbaksalyar 9 minutes ago||
Also his "Division by Zero":

https://web.archive.org/web/20111121100139/http://www.fantas...

https://en.wikipedia.org/wiki/Division_by_Zero_(short_story)

quirino 16 minutes ago||
Ted Chiang is incredible, my favorite writer.

I also recommend "Exhalation", though that has nothing to do with AI.

GMoromisato 1 hour ago||
I liked the metaphor of a climber teleported to the top of a fog shrouded mountain. And I agree that now that the teleporter exists, we need to use it to reach more peaks and explore. There's no going back to a world where AI doesn't exist.
lumost 1 hour ago|
The issue is ownership, we have no means of distributing the knowledge from the AI or rewarding those who could help.

We are quickly moving to a world where all symbolic and numeric reasoning for economic purposes is performed by AI.

GMoromisato 19 minutes ago||
Agreed! Specifically, compensation (monetary and reputational) for professional mathematicians was bundled into theorem proving--essentially, climbing the mountain. Now that a teleporter exists, we need to unbundle compensation.

I don't know what that means in practical terms, but I agree that's the issue.

yewenjie 52 minutes ago|
> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.

^^ half of the comments on this thread

ssfdg 43 minutes ago||
Also a ton of comments in this thread: breathless frothing hype declaring mathematics is over and assuming these proofs are exactly what they claim they are at face value, giving the company with a vested interest in everyone unquestioningly believing this is all real every conceivable benefit of the doubt
azan_ 41 minutes ago||
Didn't top math researchers call AI progress absolutely real and dangerous for math? It's not just HN commenters that are impressed!
ssfdg 32 minutes ago||
By all accounts the "dangerous for math" claims seem to be primarily around flooding the field with complicated impossible-to-understand proofs that according to recent research may or may not be correct depending on what's going on with the Lean implementation.

It's looking to me like it's more of a slop PR problem than it is that these things are genius at math and will displace mathematicians. I am happy to be wrong but I strongly suspect the next few weeks to months will result in more and more of this work being exposed as slop.

These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?

12kajh 48 minutes ago|||
If you haven't made progress in Quantum Computing in the last 10 years, lecturing others can become a popular pastime.
azan_ 42 minutes ago|||
You know, if you attack ad personam you've got to be ready that someone will do same against you - what progress did you make in the last 10 years (or in your entire life for that matter)?
jamiek88 34 minutes ago|||
He was created 14 minutes ago, give him a break!
moffkalast 33 minutes ago|||
Trust me bro, just 100 more cubits, I swear we'll break everyone's encryption and cause the downfall of society, please bro just one more grant, It'll be stable this time :'(
Fraterkes 31 minutes ago|||
Having stuff explained to you in patronizing tones? How horrible Scott!
Rover222 37 minutes ago|||
more like 3/4 of the comments but yea
More comments...