> Basically the paper is so horribly written that it’s impossible to read it without AI help
That's interesting and haven't seen this in all the coverage of this event.
It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.
This looks like an AI IPO PR powerplay, because at this point the proofs haven't been checked and it may not be possible for a human to check them - because proofs should be clear, not horribly written and noisy.
The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
You want the path through the maze to be as short as possible and the map to be as clear as possible.
This sounds like the opposite. There may be a genuine path through the maze, but if it's too convoluted and takes too long it will be impossible to confirm.
I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
I suspect that's possible without tripping over the halting problem. (But I can't prove it.)
In other domains I have seen first hand overwhelming evidence of how things that cause the AI to make mistakes also cause humans to make the same mistakes.
I wonder if the proofs being produced that are hard for humans to interpret are also hard for other LLMs to interpret.
In other words, I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.
It kind of an explicit example of how the LLMs can be materially less intelligent than people, but still be more productive through scaling, and yet they also can't replace people because they are a categorically different kind of intelligence. It's like all the AI debates compressed into one example showing countwr-intuitive answers.
Chow, T. Y. (2008). A beginner’s guide to forcing (arXiv:0712.1320). arXiv. https://doi.org/10.48550/arXiv.0712.1320
> “All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves. The proofs should be “natural” in Donald Newman’s sense [13]:
> This term . . . is introduced to mean not having any ad hoc constructions or brilliancies. A “natural” proof, then, is one which proves itself, one available to the “common mathematician in the streets.””
Also from what I can tell from the few fields I understand, the proofs aren't that long or complicated they are just terribly written.
https://en.wikipedia.org/wiki/Inter-universal_Teichmüller_th... seems like a counterpoint, but IANAM. (I am likely cherrypicking the far end of the bell curve re: straightforward here)
> This looks like an AI IPO PR powerplay,
Interestingly, the post has actually also an argument for this:
> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
> So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.”
Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. Mathematicians will get there.
Why would anyone believe this (also) is not simply example N+1 of this is the worst it will ever be, as opposed to recognizing this as what will almost certainly prove to be an awkward moment, soon to be replaced by another order of magnitude of cleaner, clearer, more intelligible, etc.?
Ximm's Law: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon.
Already the unit distance proof was substantially human-edited (per Thomas Bloom). Then with the ten problems from Astra you started getting the citation issues. Then Navier-Stokes was a rushed 160 pages with barely any citations, and some of the related papers were called (by their "authors") the ugliest mess they've ever seen.
And now here we are. At least it seems that mathematical ability and communication with a mathematical audience are independent skills, and progress in the first does not imply the second.
This doesn't surprise me much, given two analogies: 1) many smart people are nonetheless horrible lecturers. (You can't quite get the opposite extreme, since to explain math well you have to be able to do it.) 2) AI writing in general hasn't improved. The models have annoying verbal tics ("honestly") and have no sense of which part of what they say is obvious and which is relevant.
> mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!
My 8yo talks exactly like that. I could totally imagine him saying this, the same way, at the dining room table.
I asked ChatGPT "pretend you're an 8/9 year old today. how would you insult your mom about having her job be replaced by an AI?", and the responses it offered were:
> “Mom, AI took your job because apparently even robots were like, ‘Yeah… we can do this better.’”
> “Mom, congratulations! You got replaced by a computer. Even Siri has a job now and you don’t!”
> “Mom, AI took your job? Dang. I guess even a robot looked at your work and said, ‘I got this.’”
> “Don’t worry, Mom. You can still be useful… like teaching the AI how to make my lunch.”
All of these seem to have a vaguely Millennial flavor, aside from being pretty awkward and mechanical roasts. Trust the children and linguistic drift to be the best AI detector.
Many math problems are practically useless if you only care about the answer, the millennium prize about the Navier-Stokes equation is such a problem. The solution makes no physical sense, real life fluids don't follow the Navier-Stokes equations in such extreme conditions. But in the process of finding the solution, we may get insight into what will end up being really useful. The big mess that OpenAI produced is the solution no one really cared about, but it didn't deliver much of what people actually wanted.
One reason it is sometimes seen negatively despite being at least something is that it broke the incentive. Without the million dollar prize and with only the privilege of being second, people are much less likely to go for the insightful solution.
And going from zero proofs to one proof (even a sloppy one) is a big deal regardless of whether it was written by AI or a human.
Also lean proofs are notoriously tedious and slow to write so this level of output is very likely to be from LLMs.
Has anyone verified any of the proofs produced by OpenAI or is everyone just assuming that it just be true because the Lean code checks out? Couldn’t the Lean code just be formulated incorrectly?
For example, the statement of e.g. Fermat's last theorem in Lean should be understandable to anyone who played The Natural Number Game [0] and knows a bit of mathematics and programming. For the proof, you trust the compiler.
The statement of other theorems can be much more delicate, and the Lean formalization may require an extensive introductory section which will need to be carefully checked.
Then there are the cases where no Lean formalization is currently available, and all we have right now is an often impenetrable pdf in the OpenAI repo. I would not at all be surprised if some of those claims contained logical gaps.
Time will surely tell, but there are certainly doubts and lots people are very busy checking these results.
Highly recommend reading it. Very prescient for something written 26 years ago.
https://gwern.net/doc/fiction/science-fiction/2000-chiang.pd...
I also recommend "Exhalation", though that has nothing to do with AI.
We are quickly moving to a world where all symbolic and numeric reasoning for economic purposes is performed by AI.
I don't know what that means in practical terms, but I agree that's the issue.
^^ half of the comments on this thread
It's looking to me like it's more of a slop PR problem than it is that these things are genius at math and will displace mathematicians. I am happy to be wrong but I strongly suspect the next few weeks to months will result in more and more of this work being exposed as slop.
These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?