Posted by nill0 6 hours ago
The formalization went through, but there were _several_ mistakes in the original paper that it uncovered, from type setting errors to (many) formulas that quantified over all resources as printed, but actually applied to only arising resources in the calculus..
So the formalization did give me a formally verified borrow checker that I could use to build a programming language on top of, but it was _not_ exactly the borrow calculus that was printed in the paper.
I expect this is the most common experience when mechanizing a printed paper. There are a lot of skipped steps and handwaving.
The scary thing is when AIs generate unreadable formal proofs and then effectively lie (or fabulate, to be polite-ish) about the natural language version of the steps. Since the natural language version is arguably the most important aspect of a solution to a flagship problem, this fabulation deflates the value of the solution while the existence of the solution discourages further work on the problem.
I think a lot of math notation isn't wrong given a context, so in theory we should be able to translate it into something formal. Maybe also generate living documents where you can e.g. write `h : some_claim := by details(by rw[nat_mul_comm]; ...)` and the renderer hides details just like you'd write "obviously" in a traditional text. If the reader wants, they could then expand the details. etc. I found that many codex-generated proofs could be improved by telling it that I want a sequence of steps
have next_step := by <I don't care>
have therefore := by <still don't care>
So that the human proof appears as the left side, and I just ignore the right side as petty details. Again, not fantastic success, but better. Otherwise it goes very... Leanish by default.Lean's VSCode plugin is I think only starting to explore the idea of a proper IDE for math. There's probably still tons of unexplored potential for like that fused with Matlab or whatever.
We should be very careful about relinquishing sorting through such details to AI.
I've been criticized for doing this, but to me it emphasizes how much attention goes to the hot, wrong papers.
The idea that an AI company is beyond peer review is harmful.
that's not the claim. the formal statement of the problem for the NS proof was written by humans not autoformalized.
https://github.com/google-deepmind/formal-conjectures/blob/8...
Maybe read the comment before replying, at a minimum.
i havent seen this sentiment expressed anywhere, have you?
isn't this comment chain on a submission about openai's claims being reviewed?
A good review does not merely check the correctness of logical arguments, it gives suggestions for the exposition, citing the correct references, putting everything in the right context, etc.
Prestige to the reviewed, not to the reviewer.
Good, I just wanted to point out that peer review isn't primarily an arbitrage of truth, it is also to make sure the exposition is nice to read. When you get a reviewer who actually cares, you receive lots of feedback that isn't related to the correctness of Lemma 3.14.15 and stuff like that.
No. The way to build confidence that your software is well made, you do a proper external security audit and obtain the requisite certificate from a proper auditing firm.
It's also incorrect to think peer review in mathematics is low quality (like it is in some other fields). Certainly, when major results are in place, editors ensure that high quality peer reviewers are recruited and do their job properly. Like all human processes this fails sometimes, but not enough to not do it.
which specific openai statements does this part of your analogy map to?
in the "sharing ai progress in mathematics" blog, openai simply says "results", and never once claims that all of them are unquestionably true. instead, they state they want to evaluate the results. their github states that the results are "different stages of verification" and also says "Some of the unformalized results could have issues"
that is the opposite of "claiming [...] their software is secure", to use your analogy.
The proof was released in the spirit of being first at all costs without any attempt to clean it up. I doubt that OpenAI mathematicians could give a coherent talk about it, certainly not using a blackboard.
are people not reviewing openai claims right now?
openai themselves specifically call out that there may be issues with their results. journalists and laypeople just happen to skip that part, like they do with ~all physics, health, astronomy, etc results.
Peer review is a proxy for correctness.
Peer review journal is a proxy for quality peer review, or at least it was, once upon a time.
This is far more efficient and they’re telling the academic industry to grow up
Sister comments are saying that academics dont like the Lean programming language and see a lack of human language described proof. Doesn’t sound like something I should care about but I’m watching for a better human language description of the problem as this discussion evolves
yet i have never seen anyone say "the idea that physicists are beyond peer review is harmful" because some mainstream news articles published a piece about dark energy or whatever.
Also, it doesn't seem that they are questioning the truthfulness of either proof, just that they are different?
Actually, they are questioning whether the natural language description of the proof is either not faithful to the formal proof, or simply wrong, or both.
https://terrytao.wordpress.com/2026/10/04/on-classical-solut...
Humans will have to wade through mountains of slop to decipher the argument. Alternatively, they could just ignore it like Mochizuki's ABC proof prior to the Scholze/Stix refutation.
The point of the article is that natural language is not these things.
If you understand the Lean, then you can create a NL proof. The LLM clearly doesn't understand the Lean code it produced.
This is what I've been wondering about with LLM proofs. Math is logical, but mathematical writing is still natural language: symbols get overloaded, conventions go unstated, and a lot rides on context. So a model can translate a statement into a formal system and prove it, and the proof can check out, while the statement it proved isn't quite the one the mathematician meant. I read this article as a caution that some of the LLM proofs announced so far may not hold up once a human checks what was actually proved. Is that a fair reading?
Edit out vulgarity
It's not the form language that is the real problem here. It's the ambiguity on the other side and the extreme difficulty of doing a useful and accurate translation.
> gotcha bitch!
You may have misdiagnosed the problem.
Before it was dropping databases or deleting repositories. Now it’s subtly changing the meaning of math problems to get a correct but irrelevant answer.
> The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra.
So it suggests that the formalization/verification step may have fixed some issues in the natural language proof, and either such differences were never noticed or the corrections weren't ported back to the NLP.
Oh, well. I suppose I should avoid getting involved in these AI threads, but now it’s about half the forum.