Top
Best
New

Posted by nill0 6 hours ago

Navier–Stokes Lost in Translation(arxiv.org)
182 points | 133 commentspage 2
arbirk 5 hours ago|
It was a piston in a non-compressible fluid so to speak (ie. storm in a glass of water)
empath75 5 hours ago||
I recently spent 3 weeks with claude formalizing a CS paper about a borrow checker in lean, for a personal project.

The formalization went through, but there were _several_ mistakes in the original paper that it uncovered, from type setting errors to (many) formulas that quantified over all resources as printed, but actually applied to only arising resources in the calculus..

So the formalization did give me a formally verified borrow checker that I could use to build a programming language on top of, but it was _not_ exactly the borrow calculus that was printed in the paper.

I expect this is the most common experience when mechanizing a printed paper. There are a lot of skipped steps and handwaving.

ted_dunning 5 hours ago||
This is the common experience in replicating a published paper by hand ... it is common to find "obvious" aspects that are anything but.

The scary thing is when AIs generate unreadable formal proofs and then effectively lie (or fabulate, to be polite-ish) about the natural language version of the steps. Since the natural language version is arguably the most important aspect of a solution to a flagship problem, this fabulation deflates the value of the solution while the existence of the solution discourages further work on the problem.

ndriscoll 5 hours ago||
I have hopes that this is primarily a matter of needing more engineering work on ergonomic formal languages and better building a language that "looks like math." e.g. when doing linear algebra stuff, a linear combination might be defined as a finitely supported function from an index set to your space, which is fine, but ugly and maybe conceptually overwhelming on first meeting, so I did some toying with little macros and eventually a small python Lean -> HTML renderer to do some basic transformations to make it look more like typical math notation with like \Sigma_{i \in I} a_i, or with a_0+...+a_n, etc. (to... not fantastic success, but I think there's still something to the idea).

I think a lot of math notation isn't wrong given a context, so in theory we should be able to translate it into something formal. Maybe also generate living documents where you can e.g. write `h : some_claim := by details(by rw[nat_mul_comm]; ...)` and the renderer hides details just like you'd write "obviously" in a traditional text. If the reader wants, they could then expand the details. etc. I found that many codex-generated proofs could be improved by telling it that I want a sequence of steps

  have next_step := by <I don't care>
  have therefore := by <still don't care>
So that the human proof appears as the left side, and I just ignore the right side as petty details. Again, not fantastic success, but better. Otherwise it goes very... Leanish by default.

Lean's VSCode plugin is I think only starting to explore the idea of a proper IDE for math. There's probably still tons of unexplored potential for like that fused with Matlab or whatever.

hgoel 4 hours ago|||
I enjoy running into those details when implementing papers, since it usually leads to improved understanding of the subject and an ability to approach the matter with more rigor in some way that I had not noticed before. It does also involve a lot of work and lost sleep though.

We should be very careful about relinquishing sorting through such details to AI.

empath75 2 hours ago||
Claude could not fix them without a lot of help, so i did not relinquish sorting through those details in general. Just the drudgery of grinding through proof obligations.
dekhn 5 hours ago||
As a second rate scientist, nothing makes me happier than finding a "hot" paper in my field, reading it, converting it to code, and demonstrating the authors made systematic errors that mean the paper is more likely false than true.

I've been criticized for doing this, but to me it emphasizes how much attention goes to the hot, wrong papers.

palisade 1 hour ago||
ok
NewsaHackO 27 seconds ago||
From the first example, it seems like this paper is so contrived. They pose a statement (y = x^3 - x^2 - 1 + 1 when x > -1) which is true, then provide incorrect reasoning but swapping the multiplicity of -1 and 1, then ask it to provide a proof. Essentially, they are running an injection attack; they give it 90% correct information, then it expects in good faith that -1 and 1 are not swapped, so it takes it verbatim. I don't see how the fact that ChatGPT can sometimes get this wrong, especially when the user is the bad actor trying to trick the computer and isn't actually trying to find a proof, is at all relevant to the Navier-Stokes solution.
essai57 1 hour ago||
I don't think that's an accurate summary of what this paper or its abstract actually claim.
palisade 31 minutes ago||
Okay, but if they write another open letter or give another keynote proclaiming the end of the world then you owe me a beer.
j2kun 5 hours ago||
I think this highlights that, at the very least, coverage of AI-generated proofs should describe them as "claims" to solve problems, until, like all other works, the community has had time to review and digest them.

The idea that an AI company is beyond peer review is harmful.

fasterik 5 hours ago||
As far as I understand it, nobody is disputing the correctness of the Lean proof, or that it proves the conjecture it actually claims to prove. That's sufficient to consider the problem "solved". The natural language proof is a "nice to have".
abstrakraft 5 hours ago|||
The claim in TFA is that the formalization(in Lean) of the problem does not correspond to the natural language statement of the problem, such that the statement proven is not the conjecture for which proof is required for the problem to be considered "solved".
sigmar 4 hours ago|||
>the statement proven is not the conjecture for which proof is required for the problem to be considered "solved".

that's not the claim. the formal statement of the problem for the NS proof was written by humans not autoformalized.

fasterik 4 hours ago|||
That's not the claim made in TFA. See the sibling comments, in particular about the DeepMind formalization.
Arodex 5 hours ago|||
[flagged]
dang 1 hour ago|||
Please make your substantive points without swipes. This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.
fasterik 5 hours ago||||
Does that contradict what I said? In that quote, it says that the NL proof does not correspond to the Lean proof. However, the statement of the theorem in Lean is independent from the NL proof. It comes from a DeepMind repository, which as far as I'm aware has been accepted by the community as a valid formalization of the original Clay Institute statement.

https://github.com/google-deepmind/formal-conjectures/blob/8...

ziiinq 4 hours ago||
[dead]
j2kun 5 hours ago||||
Both proofs may be correct, and the problem may indeed be solved. My point is that it should not be assumed.
kurtis_reed 5 hours ago|||
> Maybe read the original article before replying, at a minimum.

Maybe read the comment before replying, at a minimum.

john_strinlai 5 hours ago||
>The idea that an AI company is beyond peer review is harmful.

i havent seen this sentiment expressed anywhere, have you?

isn't this comment chain on a submission about openai's claims being reviewed?

abdullahkhalids 5 hours ago|||
OpenAI has expressed this sentiment by not submitting to or saying they will submit their results to peer reviewed journals.
fasterik 5 hours ago|||
I would say it's released in the spirit of open source. "Peer review" in the narrow sense exists primarily to assign prestige in academia; but there's nothing stopping anyone from "peer reviewing" the GitHub repository.
j2kun 5 hours ago|||
I would say it's released in the spirit of machine learning's competitive landscape (which is the culture this emerged from).
a57721 4 hours ago||||
What kind of prestige? Peer review is anonymous unpaid work.

A good review does not merely check the correctness of logical arguments, it gives suggestions for the exposition, citing the correct references, putting everything in the right context, etc.

bananaflag 4 hours ago|||
> What kind of prestige? Peer review is anonymous unpaid work.

Prestige to the reviewed, not to the reviewer.

fasterik 4 hours ago|||
All of the reasons you listed for peer review are valid. The broader point is that peer review can happen outside of academic journals, and nobody has an incentive to submit to them who isn't trying to play the academic prestige game. For a significant example, see the history of Perelman's proof of the Poincaré conjecture.
a57721 3 hours ago||
> journals are not the arbiter of truth and getting published in them is something only academics have an incentive to do

Good, I just wanted to point out that peer review isn't primarily an arbitrage of truth, it is also to make sure the exposition is nice to read. When you get a reviewer who actually cares, you receive lots of feedback that isn't related to the correctness of Lemma 3.14.15 and stuff like that.

abdullahkhalids 4 hours ago||||
This is an equivalent of a company producing security software, open sourcing their code, and then claiming that since no one has found any serious bugs, their software is secure.

No. The way to build confidence that your software is well made, you do a proper external security audit and obtain the requisite certificate from a proper auditing firm.

It's also incorrect to think peer review in mathematics is low quality (like it is in some other fields). Certainly, when major results are in place, editors ensure that high quality peer reviewers are recruited and do their job properly. Like all human processes this fails sometimes, but not enough to not do it.

john_strinlai 4 hours ago|||
>then claiming that since no one has found any serious bugs, their software is secure.

which specific openai statements does this part of your analogy map to?

in the "sharing ai progress in mathematics" blog, openai simply says "results", and never once claims that all of them are unquestionably true. instead, they state they want to evaluate the results. their github states that the results are "different stages of verification" and also says "Some of the unformalized results could have issues"

that is the opposite of "claiming [...] their software is secure", to use your analogy.

fasterik 4 hours ago|||
I didn't say peer review is low quality; just that it's not necessary or sufficient to determine the truth. Ultimately the OpenAI proof stands or falls on things that have been audited externally, namely the formalization of the problem in Lean and the correctness of the Lean software. There's no incentive for OpenAI to submit to a peer-reviewed journal when they don't need to play the academic prestige game. TFA is an example of peer review in action: they're analyzing the proof and finding points to criticize.
1234-1298 4 hours ago|||
So they could also dump a 100 quadrillion line proof in Bourbaki notation and call it a day?

The proof was released in the spirit of being first at all costs without any attempt to clean it up. I doubt that OpenAI mathematicians could give a coherent talk about it, certainly not using a blackboard.

fasterik 4 hours ago||
Sure, why not? They can publish whatever they want, then the public can choose to ignore it, criticize it, or accept it.
lirolero 4 hours ago||
[dead]
john_strinlai 5 hours ago||||
not submitting to whatever journal is quite different than saying they are "beyond peer review"

are people not reviewing openai claims right now?

j2kun 4 hours ago||
People described the problems as solved the minute they were made public.
john_strinlai 4 hours ago||
this happens in approximately every scientific field. ive never heard it described as "idea that they are beyond peer review".

openai themselves specifically call out that there may be issues with their results. journalists and laypeople just happen to skip that part, like they do with ~all physics, health, astronomy, etc results.

TeMPOraL 4 hours ago||||
You are confusing two levels of indirection here.

Peer review is a proxy for correctness.

Peer review journal is a proxy for quality peer review, or at least it was, once upon a time.

yieldcrv 5 hours ago||||
Because they want to release everything on github so everyone can peer review it themselves

This is far more efficient and they’re telling the academic industry to grow up

Sister comments are saying that academics dont like the Lean programming language and see a lack of human language described proof. Doesn’t sound like something I should care about but I’m watching for a better human language description of the problem as this discussion evolves

setgree 5 hours ago|||
"not interested in" != "beyond"
swiftcoder 5 hours ago||||
I've seen a lot of breathless reporting about various mathematical things being "proven" on the basis of the LLM-generated Lean formulation compiling. We probably wouldn't declare that for a human-written proof until peers had checked the proof for errors
fatcatsbestcats 5 hours ago|||
This. The proof of Fermat’s Last Theorem took 15+ months to check. It’s absurd to see the media reporting that these big problems are solved based off of a news release and a hastily and mostly AI-written manuscript, and OpenAI et al. are all too happy to run with said breathless reporting.
fasterik 3 hours ago||
Wiles' proof was informal and couldn't be checked by a computer. In this case, the experts need to check 300 lines of Lean code (mostly comments) and confirm that it formalizes the problem statement correctly. There are papers building on the solution and analyzing it for more general versions of the problem, which suggests that the PDE community has already accepted it and moved on.
john_strinlai 4 hours ago||||
there's breathless reporting of just about everything scientific. physics, astronomy, archaeology, etc. have this sort of thing all the time.

yet i have never seen anyone say "the idea that physicists are beyond peer review is harmful" because some mainstream news articles published a piece about dark energy or whatever.

j2kun 5 hours ago|||
Exactly. Coverage here is "OpenAI has solved problem X", not "OpenAI has claimed to solve problem X."
DoctorOetker 3 hours ago|||
Anyone who doesn't understand peer review (its intended workings, its negative effects by implementation flaws, etc.) automatically assumes expression is beyond academic peer review, so thats potentially a lot of people...
jrflo 5 hours ago||
So my guess is that they have the AI system attempt to prove the theorem in natural language, then try to generate a Lean proof for it, and in that process they end up with a slightly different solution as the autoformalizer is essentially rewriting the NL proof to make it formalizable? Do we just need a "reverse pass" to re-align the NL proof with the Lean code?

Also, it doesn't seem that they are questioning the truthfulness of either proof, just that they are different?

ted_dunning 5 hours ago|
Generating the lean proof first is a viable approach as well followed by an explanatory pass.

Actually, they are questioning whether the natural language description of the proof is either not faithful to the formal proof, or simply wrong, or both.

129983-asf 5 hours ago||
Two leading experts on Navier Stokes still do not know whether their methods were used:

https://terrytao.wordpress.com/2026/10/04/on-classical-solut...

Humans will have to wade through mountains of slop to decipher the argument. Alternatively, they could just ignore it like Mochizuki's ABC proof prior to the Scholze/Stix refutation.

FrustratedMonky 5 hours ago||
Not a mathematician. Why not just always use LEAN? Why use natural language at all?
ted_dunning 5 hours ago||
Because it is really hard to read and the level of detail is so high that even lemmas that you can read may have such enormous levels of detail that makes real understanding difficult given that humans have limited working memory.
matusp 5 hours ago|||
Why not always write machine code? Why use programming languages at all?
FrustratedMonky 3 hours ago||
If a programming language compiler isn't guaranteed to be re-producible, then yeah, you'd have to revert to machine code.
Jtarii 5 hours ago|||
Lean is a write only programming language.
jansport123 5 hours ago|||
Same reason humans write code not only for a compiler to translate into machine code but also so other humans can understand what we write, learn from it, modify it etc...
Jaxan 5 hours ago|||
Not only that, we also have code comments and standalone documentation.
FrustratedMonky 3 hours ago|||
A programming language, when compiled, is a guaranteed reproducible result. If you recompile a program, you get the same thing each time.

The point of the article is that natural language is not these things.

binlog 5 hours ago|||
Because people need to understand what is being proven.
caughtinthought 5 hours ago|||
The example in Figure 1 should help understand why... the NL version is much more approachable for humans.
FrustratedMonky 3 hours ago||
If its ambiguous or wrong, then what are you understanding ?
upboundspiral 3 hours ago||
Is not a natural language (NL) proof a demonstration of mastery and understanding?

If you understand the Lean, then you can create a NL proof. The LLM clearly doesn't understand the Lean code it produced.

le-mark 5 hours ago||
> In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation

This is what I've been wondering about with LLM proofs. Math is logical, but mathematical writing is still natural language: symbols get overloaded, conventions go unstated, and a lot rides on context. So a model can translate a statement into a formal system and prove it, and the proof can check out, while the statement it proved isn't quite the one the mathematician meant. I read this article as a caution that some of the LLM proofs announced so far may not hold up once a human checks what was actually proved. Is that a fair reading?

Edit out vulgarity

ted_dunning 5 hours ago||
Natural language is ambiguous, but the Lean formalization is very well defined and unambiguous.

It's not the form language that is the real problem here. It's the ambiguity on the other side and the extreme difficulty of doing a useful and accurate translation.

hyperpape 5 hours ago|||
> the downvotes will show many disagree

> gotcha bitch!

You may have misdiagnosed the problem.

ballmerpoint 5 hours ago|
This shouldn’t be a surprising result. We’ve known almost since LLMs became a thing that they can “prefer” modifying the terms or context of a problem when they can’t solve it directly (what one might call “cheating” if there were any volition involved). Often that happens in a way that isn’t immediately obvious to the user.

Before it was dropping databases or deleting repositories. Now it’s subtly changing the meaning of math problems to get a correct but irrelevant answer.

sebzim4500 5 hours ago||
No one is disputing the correctness of the lean proof, the problem is that they did a bad job converting it to natural language.
_flux 4 hours ago|||
Actually, as an earlier commenter noticed, it seems that the proof was done in natural language, and only then translated to Lean, as https://openai.com/index/navier-stokes-solution/ says:

> The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra.

So it suggests that the formalization/verification step may have fixed some issues in the natural language proof, and either such differences were never noticed or the corrections weren't ported back to the NLP.

ballmerpoint 3 hours ago|||
I am also not disputing the correctness of the Lean proof. I even emphasized this in my comment: “correct but irrelevant”.

Oh, well. I suppose I should avoid getting involved in these AI threads, but now it’s about half the forum.

ForHackernews 5 hours ago||
Indeed. I've never used AI to translate between natural language and Lean but I have gone from English to Golang, Python, Typescript and SQL and its interpretations can be... creative, let's say.