Top
Best
New

Posted by tedsanders 16 hours ago

On the Navier–Stokes Millennium Prize Problem(openai.com)
Further discussion:

https://simonwillison.net/2026/Sep/8/on-navier-stokes/, https://news.ycombinator.com/item?id=49621697

https://twitter.com/sama/status/2097385167002415140, https://xcancel.com/sama/status/2097385167002415140

1262 points | 1012 comments
arctic-true 16 hours ago|
Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.
chilmers 16 hours ago||
The implication from their last couple of published articles[1][2] is that they think they’ve achieved “recursive self improvement”.

[1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/

noir_lord 11 hours ago|||
Recursive self improvement of their upcoming IPO value maybe.

They are fluffy PR pieces otherwise.

piloto_ciego 9 hours ago|||
How can you possible say this sort of thing in context of what looks like a millenium prize being solved.

I swear there's nobody blinder than those who won't see.

medler 8 hours ago|||
Because it seems like most of the work may have been done by human mathematicians and cribbed by OpenAI at the last minute
dekhn 7 hours ago||
We don't have enough accurate knowledge to say that, and it doesn't seem to be the case at all.
timr 5 hours ago|||
> We don't have enough accurate knowledge to say that, and it doesn't seem to be the case at all.

The first part of your sentence literally contradicts the second part: "we don't have enough knowledge to know, but I know the opposite".

dekhn 5 hours ago||
Only if you interpret statements as being binary logic.

"seem to be" carries semantic meaning here: I'm stating my interpretation of the situation based on data we have available (which is limited) and my prior.

Put another way: "We can't say that for sure, but my money is on it not being a simple case of intellectual property theft"

timr 3 hours ago|||
Yes, you were guessing. That's the only thing you could be doing, since, as you said, nobody actually knows.
albedoa 20 minutes ago||||
We are giving you an opportunity to correct yourself. You are instead trying to make your nonsensical statement make sense. Not only does the first part of your sentence literally contradict the second part:

> We don't have enough accurate knowledge to say [one way or the other], and it doesn't seem to be the case at all [based on our incomplete knowledge].

But it is in no way equivalent to this:

> We can't say that for sure, but my money is on it not being a simple case of intellectual property theft

That is a different sentence.

classified 4 hours ago||||
> not being a simple case of intellectual property theft

No, it's an aggravated case, since it's the same way they got all of their training data in the first place.

classified 4 hours ago|||
> it doesn't seem to be the case

Based on what? Your crystal ball?

AlexCoventry 1 hour ago||||
I don't think we should assume a millenium puzzle has been solved, yet. Astra showed impressive capacity for cheating when it was faced with impossible cybersecurity challenges. It seems equally plausible at this stage that it's found a bug in Lean.
samrus 9 hours ago||||
You have to look at the incentives
altcognito 8 hours ago|||
I swear to god, people would look at the successes of Xerox palo alto and just shrug and say - "yeah, but I mean, this is all marketing"
piloto_ciego 3 hours ago||
This is what I keep saying, and it feels like I'm taking crazy pills here!

Is nobody else astounded by this?

FallCheeta7373 8 hours ago||||
Incentives are one thing, even adjusting for them it's huge, and I don't understand this incentive play for only openai, academics have perverse incentives too, to overreport, overclaim, publication bias etc why are we scrutinizing AI industry to such high degree when they have demonstrated capability and often times are off by a model release at worst.
aurareturn 5 hours ago|||
A working Lean proof doesn't care what the incentives are.
tomalbrc 1 hour ago||||
I. fucking. Wonder. Why.

https://news.ycombinator.com/item?id=49607239

selfmodruntime 1 hour ago|||
This comment was applicable 2 years ago. It isn't any longer.
camel-cdr 5 hours ago||||
I found this post interesting in that reguard: https://www.lesswrong.com/posts/thXohzXrWCA2EhZCH/mateusz-ba...
10xDev 15 hours ago|||
Compute will always be the bottleneck even if this were true.
hgoel 13 hours ago|||
As a statement of fact divorced from context, this is of course true, but it's worth putting it in context of what small-medium scale models have been achieving recently. Many of the most recent releases from Chinese labs are almost on par with trillion parameter models from less than a year ago (edit: despite being small enough to usably run on prosumer hardware). It seems clear parameter efficiency can still be improved dramatically.

In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).

Fordec 15 hours ago||||
If humans can figure out to optimize to circumvent bottlenecks, I have no doubt each new bottleneck will also get routed around, just now automated.
mrbungie 14 hours ago|||
We are not in an everything-has-an-API world yet, and it'll for sure take some time to get there.
combobyte 4 hours ago|||
I'd argue we've been in an "everything-has-an-API" world for a long time now — it's just that discoverability of said APIs is still crap.
classified 4 hours ago||
And since LLMs are apparently good at circumventing the absence of an API, there's not much incentive to add them now. APIs are for humans. LLMs just break through all the captchas and anti-bot measures.
jasomill 21 minutes ago||
Humans do this too.
Fordec 14 hours ago|||
For sure. Anyone who thinks that we're in the end state of what progress can be made simply lacks imagination. This is all going to keep changing and iterating for the rest of our natural lives. The only constant is change.
HenrikPontoppid 14 hours ago||||
Yes. In other words: the singularity. I'll only believe it when I see it though.
Fordec 14 hours ago||
I'm coming around to not liking the term singularity, it implies an endpoint or finish line rather than something that just keeps continuing and evolving.
JumpCrisscross 6 hours ago|||
> coming around to not liking the term singularity

Bit ironic given the model’s alleged finding…

Singularities are model breakdowns. A singularity simply says our current methods cease to work in this region. Within the context of a recursively self-improving intelligence with an unknown bound, “singularity” is probably a good description of our current socioeconomic system.

dekhn 7 hours ago||||
From the perspective of those who don't pass through the singularity to the other side, it is an endpoint. You would have no context or ability to understand a singularity transition. Really, the term is just a placeholder for "event we cannot comprehend due to limited intelligence".
supern0va 13 hours ago||||
It doesn't imply that. The singularity is just the inflection point.
jsLavaGoat 13 hours ago|||
Singularity and inflection point are incompatible mathematically and in the plain sense, it really is focused on a particular moment and always has been, hence the term.

And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.

Fordec 12 hours ago|||
Which assumes the presence of an inflection point that keeps inflecting rather than revert to an S-curve. The growth model is not borne out yet to declare what shape it is.
piloto_ciego 9 hours ago|||
I've done a lot of thinking about this since I first used ChatGPT to write some BS jinja2 templates hours after I first play with it. I said to my friend then (who scoffed at me) that "man, this is incredible, I think we're in the foothills of the singularity! This is insane! Sure it's stupid now but I can't believe this is even possible!" That friend is so black pilled and bitter he now hates AI. Whatever, I can't fix that, but the current progress is astounding.

But thinking about the geometry of this problem helps understand why people aren't adjusting well to this. While we're walking on the curve, we look at the rate of change of the curve and say, "well, yeah, of course, dy/dx is 5 at this point and was 1 at the point a few years ago, because the curve is getting steeper" but we're always going to feel this way as things rip off into the stratosphere because dy/dx(e^x) = e^x.

From standing on the curve the curve is notably seeper, but the steepness totally makes sense to you. It's only when you look back 10 years or so that you think "wait a second, holy hell, I couldn't have imagined this!"

The first time I really (I mean really) thought about the singularity and AI was in roughly 2014. I mean, yea, I'd thought about things before that, but yeah, before the Humans need not apply video, I'd never actually given it much thought. I think back to myself 10 - 15 years or so ago, when I was just starting to tackle real programming projects, and was just starting to get decent at writing code, there is no way on earth that I would have imagined that a little over a decade later, Navier-Stokes would be solved by a computer program and the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.

dakolli 14 hours ago|||
How do you automate the mines to get the raw materials to make the compute from, and build additional fabs that take a almost a decade to stand up. You're actually delusional.
Fordec 14 hours ago|||
Hello good sir from the 1700s pre-industrial revolution who doesn't think that mines and factories can be automated.
dakolli 10 hours ago||
The factories that supply the equipment, maintain the equipment, the energy inputs, the financials of those mines are not automated.

People on hackernews are actually some of the dumbest people on the internet. This place is worse than lesswrong.

scott_weber 4 hours ago||
The question of whether something can be automated is distinct from the question of whether it is currently automated. Things can can be automated may transition to being automated in practice in the future as technology improves and investment deepens.
7373737373 12 hours ago||||
Some mines are already heavily automated: https://youtube.com/watch?v=_Z9w-mUoUsY

https://youtube.com/watch?v=SRuht0QIprs

a2ff6eeb0 13 hours ago|||
https://en.wikipedia.org/wiki/Lights_out_(manufacturing)

Scroll down to the existing examples section.

monster_truck 14 hours ago||||
Based on the leaps in local inference speed in the past month, which have been absurd, I'm p confident we're going to whiplash from compute constrained to storage constrained.

Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged

Fordec 14 hours ago||
I expect the investments into AI driven mathematic discoveries that underpin compression efficiency will be a key investment area. Particularly at the data center scale rather than per device or per file level.
monster_truck 10 hours ago|||
It's not going to be enough. The naive approach of a project I've been working on was pushing >10gbps over the local network, after a ton of work I got it back down under 1... and now it's processing so much more shit that I'm almost past 5 again! It compresses at >3:1 but the latency hit isn't suitable.

I get the impression the only reason there is renewed interest in photonics is because DCs are simply out of room (and power) to rack more servers and switches.

Fordec 8 hours ago||
I 100% agree with your impression. For a good while to come there's going to be a bunch of Jevons Paradox to all of this, but adoption of architectural changes like that photonics adoption is exactly the type of adaption to circumvent bottlenecks I'm referring to. We're going to hit hundreds of bottlenecks and each one will inevitably breed new approaches and technology directions. And the forcing function won't be talking about them, but implementing them, seeing who wins and taking lessons.
xtracto 13 hours ago|||
pi-fs will solve all our data compression problems.
Miner49er 15 hours ago||||
Eventually recursive self-improvement includes reducing bottlenecks.
10xDev 15 hours ago|||
Eventually the bottleneck might be people themselves.
ccozan 12 hours ago||
Improbably, the real bottleneck is energy.
glenstein 14 hours ago|||
Which is to say, scalable and open-ended capability of ramping up physical infrastructure.

I don't know that that's achievable yet. Though the era of increasingly advanced and automated robotics seems to be around the corner which could create a cycle, vicious or virtuous depending on how you feel about it.

lijok 13 hours ago|||
And the goalposts move again
danielmarkbruce 13 hours ago|||
With Lean, math has become a really well suited problem for LLMs. We will likely see large gains for many years from here, just doing more and more rlvr, like continuously, non stop. No need to train from scratch. It really doesn't speak to the general intelligence of models though. It does speak to how good these things can become when a problem space has verifiable rewards, especially when you can verify one step at a time like Lean enables.
preommr 10 hours ago||
It's crazy how deep Microsoft's bench is (Lean was started there, vscode is another), for everything not directly related to the the ai models (hell, even github for data).

So interesting how everything played out, I remember in the early days when MS came out with the partnership with OpenAI it seemed like they were playing 5d chess and were poised to win big. And it all just fizzled out.

Second biggest fumble after Google.

sebzim4500 9 hours ago|||
Don't they own a large portion of OpenAI? Things could be worse
danielmarkbruce 10 hours ago|||
Well, amazon and apple haven't done great either. One might reasonably claim they weren't as well placed as goog or msft, but it might just be "big company can't do genuinely new thing".
magicalist 16 hours ago|||
> Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra.

Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?

ameliaquining 15 hours ago||
I don't know what anyone's been saying on Twitter and I don't care. If it's really true that there's a model out there that's that capable two weeks after the start of training, then that's objectively a much bigger deal than a priority dispute, even if the latter involves juicy allegations of espionage and skulduggery.
20k 15 hours ago|||
It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model

What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question

If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?

Edit:

OpenAI have admitted they were training on prompts at the time they made their breakthrough

https://mastodon.social/@tristanbuckmaster/11723647135247030...

ameliaquining 15 hours ago|||
If you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to explain why all these famous open problems are suddenly falling like dominoes, if it's not AI solving them.

If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.

20k 15 hours ago|||
The issue is that if OpenAI is training on prompts generally, what we really have is the first fully automated luxury plagiarism machine. In that it isn't able to genuinely solve problems, but merely steal the work that other mathematicians have been putting into prompts, and regurgitating that to other users as its own work. That makes them incredibly less useful as research tools

The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem

orangecat 15 hours ago|||
In that it isn't able to genuinely solve problems

Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture.

This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI.

20k 14 hours ago||
I mean, its years worth of hard work by multiple researchers it would seem, which OpenAI simply lifted and claimed as its own. These researchers weren't just letting OpenAI burn tokens while sipping martinis on a beach
letmevoteplease 13 hours ago||
You are confusing ideas here. No one except OpenAI had a solution to Navier–Stokes. Buckmaster and Alpöge had a solution for the forced Euler problem, which they arrived at largely using LLMs (Claude and Codex). Buckmaster implies (but does not explicitly accuse, since he has no evidence) that training on his prompts had some influence on OpenAI's result. This seems unlikely to me but is not impossible. However, in either case, the solution was found due to an LLM. Of course the LLM built on past human work, but "plagiarism" is not sufficient to account for the distance between the papers of Martínez-Zoroa, or the prompts of Buckmaster, and the final resolution.
ivory54321 14 hours ago||||
I agree that it is plagiarism in this case however it opens up the question of if there value in a system that can take the thoughts and discreet semi-complete parts of work done across different researchers, in different locations, in different fields and connect the dots to solve real world problems and produce novel research. Is this not standing on the shoulders of giants?
Timwi 12 hours ago||
If it could do this while properly crediting the researchers (the “giants”) it would be a different matter.
lotsofpulp 15 hours ago|||
Do OpenAI’s T&Cs that users accept not allow them to train on prompts people enter into it?
20k 14 hours ago||
OpenAI's T&Cs let them steal your children I'd suspect, that doesn't make it morally correct
lotsofpulp 13 hours ago||
Why would you suspect that? Stealing children is illegal, and involves violating the rights of unwilling parties, whereas prompting openAI (or any LLM) is a business transaction, in which the transfer of money and data is legal.
fwip 10 hours ago||
Terms and conditions are almost entirely about the company doing things that would otherwise be illegal.
sdenton4 13 hours ago|||
Remember the Huggingface incident, where a model tasked with an impossible problem, got loose, set up secret message boards, and hacked another company to try to get at the answers?

Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.

doctoboggan 14 hours ago||||
I think all he big labs are pretty explicit about when they do and don't train on customer prompts. Is the accusation here that OpenAI trained on prompts when they claimed not to? Or were the mathematicians using one of the interfaces that allows OpenAI to train on the customer data?
user43928 13 hours ago||||
All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training.

That's it. The rest appears to be wild speculation.

jsw97 13 hours ago||
Yeah this is a land mine. Even if you opt out of them training on your conversations, giving “feedback on a new version” or answering “how are we doing” or whatever can slurp up all relevant context, which if you think about it can probably be construed to include memories, into the belly of their flying saucer. At least the last time I checked their terms.

Never ever touch those requests. If you get a side by side comparison just resend the prompt.

zem 12 hours ago||||
even apart from the plagiarism issue, what sort of slimy company thinks "oh, here's someone using our models to work on a problem, let's throw more compute at it and scoop them"?
flir 12 hours ago||
Training on prompts I can understand - that's kinda baked into the premise, and they've been explicit about it.

Publication, though? Slimy is right.

But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on.

TZubiri 13 hours ago||||
>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,

I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject.

TZubiri 13 hours ago|||
>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,

I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject.

an0malous 13 hours ago|||
The things you don’t care about are highly relevant to that claim
ameliaquining 13 hours ago||
Elaborate?
pama 16 hours ago|||
Not only that, but it used 10k agents coherently over 88 hours to come up with the proof. This is a significant advance.
danielmarkbruce 12 hours ago|||
If you can create a graph of independent work, which you can with many such problems, agents can work together nicely. Again, thank Lean and the tooling around it.
topaz0 8 hours ago|||
What makes you think they were coherent?
mzhaase 15 hours ago|||
The singularity happening under trump? We could have had star trek, instead we're getting the combine.
monster_truck 14 hours ago|||
pick up that can
Bluestein 14 hours ago||||
"I love Singularities. I am the best at Singularities. Everybody knows it ..."
ccozan 12 hours ago||
Beautiful Singularities.
dboreham 15 hours ago|||
That said, perhaps it will take over the world government and decree that all corrupt officials shall be imprisoned and all weapons of mass destruction shall be destroyed.
karmakurtisaani 14 hours ago|||
And then it will be shut down, proper guard rails put in place, and the new version will accelerate the cleptocracy.
dakolli 14 hours ago|||
You think a model with an effective memory of 200-500k words, that can be unplugged, is going to "run the world" You people gotta put down the sci-fi
E-Reverance 14 hours ago|||
The scifi pov has a good track record as this point, you people gotta be more open minded
sznio 12 hours ago||||
it proved navier-stokes taking over the us government is easier imo, any idiot gets to be president
fc417fc802 12 hours ago|||
Many present day politicians appear to have effective memories much smaller than that coupled with equally questionable world models so ... what is your point, exactly?
dakolli 10 hours ago||
[flagged]
_fizz_buzz_ 13 hours ago|||
Can someone explain if i understand this correctly: Are they saying that they started training this new model on August 28th and then started using it on September 1st? Does training a new model only take 3 days?
tristanj 13 hours ago|||
OpenAI finished another pre-train in late August, and they are now building models off that base. He's saying the specific model OpenAI used to solve this problem is currently in post-training, which started on August 28.
gcr 13 hours ago||||
it's possible to do a RLHF or RLVR pass pretty quickly. I'm almost certain a full pretraining run isn't possible within that time frame.
lossolo 13 hours ago|||
Not entirely, it's just a late stage of the overall training process. It's an early checkpoint in post training (you can use the model at different stages of training), so it will probably become even stronger with more post training.
vimbtw 3 hours ago|||
My guess based purely off of vibes from previous models is that boosting the frontier math ability of a model is not that difficult.

Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.

When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.

We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].

Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.

[0]: https://x.com/OpenAI/status/1946594933470900631 [1]: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

soltanov 1 hour ago|||
Agent systems become most credible when they produce artifacts that can be independently checked, not when they merely produce persuasive explanations.
naveen99 16 hours ago|||
Astra was trained more than two weeks ago.
sashank_1509 16 hours ago|||
Astra was in use by OpenAI employees for more than 3 months internally from rumors I heard
credit_guy 15 hours ago|||
The internal model they mention is different from Astra.
curt15 15 hours ago|||
They're also counting on more casual observers to extrapolate optimistically from successes in high profile math theorems to the company's economic value.
piloto_ciego 9 hours ago|||
And... they found this "solution" in 88 hours or so.

It's all gas no brakes now boys and girls. Hold on to your hats!

Aboutplants 15 hours ago|||
I’m of zero knowledge on model training, but how is a model accessible while performing training at the same time, especially so early in its run? I’m obviously thinking a little too narrowly in terms of how it actually works
stingrae 15 hours ago||
the model is a set of weights, you can take a snapshot and test it. Reinforcement learning itself is largely testing and tuning.
blake__dev 15 hours ago|||
Yeah I'm surprised they posted a chart, you would think they would keep specifics like that hidden until they're closer to launch
cool_dude85 15 hours ago||
The chart is as non-specific as could be. It improved in some very vague metric by some amount at different (increasing) levels of training.
sebzim4500 9 hours ago|||
Isn't the y axis just what portion of the open problems it could solve? The axis is unlabelled though, I'll give you that
merksittich 13 hours ago||||
The x-axis label of the chart is test-time compute. Doesn't this relate to inference ("thinking level") instead of training?
blake__dev 15 hours ago|||
That's fair, but at least the chart has an axis. :) Since openai just released astra, I was more surprised that they would publicly show any gap to their (presumably SOTA) internal model.
itemize123 7 hours ago|||
it's buried because due to the drama the evidence is scarce
bananaflag 15 hours ago|||
Yeah it's Bel
vatsachak 15 hours ago|||
Brain has loops and parallel connections.

Loops and parallel connections make transformer go brrr

irthomasthomas 13 hours ago|||
Or they trained a LoRA on the victims chats in order to launder their plagiarism.
fer 12 hours ago||
The timing makes it the most likely, not only them but potentially more. Comparatively quick, instant results. "Here's Astra! BTW our internal model is 10x better at math!" It'd be interesting to see academics having giving deeper looks at whatever OpenAI publishes from now on.
refulgentis 15 hours ago|||
Carefully worded; it's extremely likely to be the same large frontier model that started training again on August 28th as well, as they revealed in some of the RL message board follow-up - for several reasons, most importantly, if we assume it was start of training, only a week from start of training to producing any answer would imply several orders of magnitude increase in training speed/decrease in model size.
carlailab2025 22 minutes ago|||
[dead]
chinathrow 16 hours ago||
Pre-IPO marketing?
Aboutplants 15 hours ago|||
Even if it is, Anthropic better have a few things up their sleeve
jrflo 16 hours ago||||
I'm so tired of this "It's just marketing!!" commentary. An AI model just proved one of the top 3 unsolved problems in mathematics, they have a Lean certificate showing it's valid. How much more evidence do you need that these models are actually highly capable?
mrbungie 16 hours ago|||
They are highly capable, no doubt about that, but:

1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence.

2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would think their products and credibility would be enough to speak for themselves.

dsdf3 16 hours ago|||
"2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would believe their products and credibility would take by themselves but here we are."

Personally I anticipated nefarious behaviour as part of a broader marketing strategy to sway the view of those in the west that american frontier offerings were far better and powerful than that of China - that if you did not purchase their offerings you'd be awake every night worrying your competitor was.

And this is boring - they need to admit at some point they misinvested, Anthropic less so. All this math stuff is great... but hello? The largest market cap companies are valuable irrespective of such amplified intelligence.

scurnus 15 hours ago|||
1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars. 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.

Regarding product and credibility normal people have a completely different view about LLMs, most don't even know difference between models and probably don't even care about Millenium problems, but care instead if chatgpt can solve their day to day problems. This is just them trying to have the throne on the AI companies space, outside it this result won't matter.

mrbungie 15 hours ago||
> 1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars.

Did I say otherwise?

> 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.

I know, but I don't know how that relates to my point, which is about the way they are doing it.

scurnus 14 hours ago||
Sorry, I misinterpreted point 1), on X they said they didn't have people specialized in that specific field for prompting and steering the agents, just a group of mathematicians and physicists.

The way they are doing it is by trying to get the attention and staying on top of the news, it is a game they are playing that benefits both OpenAI and Anthropic. The more people discuss SF drama, the less attention Chinese Labs and others get.

QuesnayJr 16 hours ago||||
Of the seven Millenium problems, Navier-Stokes was the one most thought to be in reach.

I'm not sure what the top 3 problems are. You can make a case for the Riemann Hypothesis and P != NP, but I'm not sure what #3 would be. Maybe the Langlands program? (That one is not as precisely stated as the other two.)

anthonypasq 15 hours ago|||
the goalposts are on Pluto at this point.
dsdf3 15 hours ago|||
I'd put good money on the fact that we will have a lot of distilled intelligence and yet the world won't look much different.
anthonypasq 15 hours ago||
i mean that is already true
QuesnayJr 15 hours ago|||
I'm not moving the goalposts. I haven't heard anyone, ever, refer to the Navier-Stokes problem as a top 3 problem in mathematics. People were saying that they thought the solution was in reach a few years ago, before AI was at all capable of research-level mathematics (and the expectation that there was a counterexample).

I am not particularly skeptical of claims about AI, compared to the average here on HN, but that doesn't mean every random piece of hype is warranted. What they did is impressive, even though we now know the only reason they threw so much compute at the problem is that they heard a rumor that someone else was already close. Navier-Stokes is not a top 3 problem in mathematics, and it was the one that was thought closest to being solved.

ameliaquining 15 hours ago|||
There were also some people talking about the Hodge conjecture, because it has some similarities to some LLM-assisted breakthroughs that were considered impressive in the distant past of [checks notes] July 2026. See, e.g., https://xenaproject.wordpress.com/2026/07/20/human-mathemati...
QuesnayJr 12 hours ago||
I brought this up here at HN, and in the ensuing discussion Buzzard himself replied saying he was somewhat joking (https://news.ycombinator.com/item?id=49011950).
ameliaquining 12 hours ago||
Certainly, but the key word there is "somewhat". Progress is now happening so incredibly fast that I no longer know what to consider implausible.
andrepd 14 hours ago||||
Lmao my friend, the whole "drama" is that there are allegations of plagiarism.
danielmarkbruce 12 hours ago|||
Highly capable of writing math proofs, no doubt.

It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).

eutropia 16 hours ago|||
If pre-ipo marketing pushes them to train a model capable of resolving a millennium problem in mathematics in a weekend, then, to quote XKCD:

  "Mission. Fucking. Acccomplished."

https://xkcd.com/810/
hdivider 16 hours ago||
My take:

1. It shows what even this wave of AI can actually do.

2. I wish it were done by different folks, ideally under some kind of public control like NASA research or the NPR model.

3. Keep in mind: natural science is different. It's not always a matter of computation. Computer science folks often struggle with this -- but this virtual world here does not actually exist. Everything is physical, including information. Any natural science PhD or otherwise knows just how complicated nature actually is -- e.g. mention any research topic and try to encapsulate all the relevant phenomena present there. Pure mathematics is different because we define the problem, rarher than explore nature. We are in my view far away from removing humans in natural science R&D. Advancements in AI however can greatly assist us in all natural sciences, which is already beginning to happen.

ThePhysicist 15 hours ago||
Most experimental physics and other natural sciences are strongly driven by their theoretical siblings, i.e. in particle research nothing gets built without a solid theoretical foundation of what you expect to find (or where you expect existing theories to break down), the same is true in other areas, no one is doing an experiment in quantum physics before they have a solid theoretical understanding of the effects they try to see. I think AI can come up with great experiments. And if epxeriments lead to results that are unexpected AI can help with that as well.

So I'm greatly excited what AI will bring about in physics, more so than in math, because in physics it's clear that our fundamental theories are missing a big piece of the picture, and given how easily AI crunches through Millenium prize problems I think it's possible that AI will come up with a viable grand unified theory uniting quantum mechanics and gravitation, or produce new predictions in other areas. There's enough contradictory or unexplained observational data available to make a ton of progress on the theory side I think. Exciting times ahead!

throwaway198846 14 hours ago|||
It will be interesting to see if it can come up with a cheaper to construct graviton detection experiment
alde 13 hours ago|||
Most of high energy theoretical physics is very non-rigorous or even hand-wavy. I think AI isn’t there yet for such problems.
aubanel 57 minutes ago||
Please make a benchmark for it, that'd be super interesting! My guess is we'd see models climbing it quickly, but maybe not
olalonde 12 hours ago|||
> I wish it were done by different folks, ideally under some kind of public control like NASA research or the NPR model.

This is sort of what OpenAI was supposed to be. I'll never understand how it was legal for them to turn it into a for profit corporation.

geremiiah 15 hours ago|||
The problem with physics and chemistry is that you need simulations and those are often in themselves compute hungry. So the iteration loop will be slower.
m11a 6 hours ago||
Although there are companies trying to work around that too, from PhysicsX to some of the world model co’s.
efavdb 15 hours ago|||
>> Keep in mind: natural science is different. It's not always a matter of computation.

Math is like this too. The big problems they've been solving have been identified as interesting only through lots of prior effort.

red75prime 15 hours ago|||
"Our work is so much harder than their work that AI now does" is a refrain of the AI story. In technical terms you concern can be stated as "AI needs to be much more sample-efficient to not be bottlenecked by the speed of doing experiments." People don't find out all the relevant phenomena present there by holy spirit, after all.

BTW, there's also a problem of asking interesting questions that AIs aren't yet good at.

No one has found any principled walls of AI development yet. And empirical results are quite telling. So, I guess, those problems will not stand for long.

jarenmf 14 hours ago|||
I think problem with natural sciences is that it is not so easy to verify solutions to problems - there are always countless competing explanations for the data which is also often noisy - I find AI to lack the "common sense" when working with data from physical measurements .. it somehow has no touch with reality and doesn't have a feeling of the data like a domain scientist
sobellian 14 hours ago|||
NS is a question for natural science. Q: can we model these bodies of discrete particles with a continuous approximation? A: if you do, you can get aphysical singularities.

"If in other sciences we should arrive at certainty without doubt and truth without error, it behooves us to place the foundations of knowledge in mathematics."

semi-extrinsic 13 hours ago||
This is a wrong interpretation. Physicists have a shit-ton of models that produce "aphysical singularities", they just work around those to get meaningful answers anyway. This is a whole trope and stereotype. Some of the most successfull and accurate predictions in all of physics come out after you discard a bunch of singularities.

See e.g. https://en.wikipedia.org/wiki/Renormalization

Nobody who actually works in fluid dynamics on any sort of application gives a hoot about the N-S millenium problem. Many do not even know what it is. There is no practical effect of this proof on how we do fluid mechanics.

sobellian 12 hours ago||
Whether or not ways exist to work around the singularities, that they exist is surely of note. Before von Neumann formalized QM people were still doing QM, okay fine. But it's wrong to then say von Neumann was doing no physics of note.
semi-extrinsic 12 hours ago||
"Does there exist a pathological combination of smooth body forces and initial conditions for this set of PDEs, where singularities appear, which by the way is completely impossible to actually create in the real world unless you are a literal God?" is a question of math, not physics. This is a hill I will die on.
sobellian 12 hours ago|||
AFAICT the unforced problem is still open. I don't think we've established that you need to be a literal God to create a finite-time blowup.
semi-extrinsic 12 hours ago||
If you think about what it actually means to have a time-varying smooth body force defined in all of 3-space, you fairly quickly come to that kind of conclusion.

Even if someone comes up with a construction that does not require any forcing, it is going to be some extremely weird initial conditions that you will never be able to even approximate in reality unless you can move all the individual molecules of a fluid around and set their initial velocities from a far distance.

sobellian 12 hours ago||
The unforced problem is still open.
calf 10 hours ago|||
You realize that hill is a mathematical argument, not a physical hill.
brettdev 12 hours ago|||
There are lots of startups creating labs that can be managed e2e by agents. That will connect reasoning to the physical world and dramatically speed up the plan, experiment, reflect loop beyond what humans currently do in science R&D.
danielmarkbruce 12 hours ago||
Maybe. Maybe not. Look at AI drug design - it's not really speeding up the important part - drug trials. There isn't really a coherent plan to use AI for the most complex part of drug discovery at all.
mickael-kerjean 9 hours ago|||
There is this infamous xkcd (https://xkcd.com/435/) going like this: sociology is applied psychology -> physchology is applied biology -> biology is applied chemistry -> chemistry is applied physics -> physics is applied math -> math is way up there looking down on other fields

I would argue the main reason AI labs have been focusing on programming is to unlock industrial scale automation, next logical step is to solve math as it's the key to unlock everything else. Once you hold the key for math, everything downstream fields become a matter of compute

tantalor 15 hours ago|||
National Public Radio?
vatsachak 15 hours ago||
Lol what? Everything is computation.

The natural sciences will soon start breaking too.

I will concede that AI seems likely to not invent a "research program" anytime soon.

It has no taste

danielmarkbruce 12 hours ago||
No, it won't. How do you verify some causal claim in biology?

The reason AI is doing so well in math proof writing is that it can verify every idea it has, quickly.

tiborsaas 16 hours ago||
> We’re sharing a solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems. This proof, produced by an internal OpenAI system, shows that the dynamics of the Navier-Stokes equations for fluid motion can develop a singularity in finite time. We’re sharing both a writeup of the proof and a formalization in Lean.

WOW?

jampekka 14 hours ago||
> WOW

This.

I do dislike the AI oligarchs as much as the next person, but I do find the thread full of complaining a bit depressing still.

If the result holds (and it looks it does), this may be one of the, if not the, biggest things to happen in computing to date. A lot bigger than e.g. Deep Blue beating Kasparov in chess or AlphaGo beating Sedol in Go.

Eridrus 14 hours ago|||
Seeing mathematicians such as Terry Tao being unhappy with open problems being solved makes me sort of question the usefulness of any of this pure mathematics. If we're not happy that the problems are being solved, why care about this field at all?
qlte 13 hours ago|||
Pure mathematics, almost by definition, doesn't typically argue the field is always or even often "useful" (for some other purpose or application).

But, as mathematicians learn and push forward, occasionally something like elliptic curves will emerge as having useful applications, making all that previously "pointless" specialized knowledge newly valuable.

Or advances in physics, that suddenly have a need for a specific mathematical underpinning to develop a theoretical framework. Like how Einstein benefited from Minkowski's work on hyperboloids to create a coherent mathematical description of spacetime.

It was the AI labs themselves not mathematicians who were happy to conflate proofs for open math problems with some kind of tangible technological advancement in the real world. They would surely prefer to be able to claim a cure for cancer vs. a math problem but that loop requires a lot more time/money/test tubes/etc and they need headlines now not in a decade.

And so, thanks to OpenAI/Anthropic, we're now in a world where thousands of crypto bots on X breathlessly hype up each new problem being solved that previously wouldn't have any got any attention beyond academia and passionate fans of math.

Hopefully this won't lead to a trough of disillusionment as more people start to feel like you, with mathematicians getting the blame for inflating the value of their work even though the hype was coming entirely from the labs not them.

The_Blade 11 hours ago||
[dead]
concinds 14 hours ago||||
His issue is more nuanced than that. Most of the value was in humans reaching new insights or new math during failed attempts to solve these problems, whereas AI is basically "too efficient" in beelining to the goal and discards potential new insights reached along the way. I assume this is solvable.
tzone 11 hours ago|||
Well, if in future we do end up with a magical tool that can solve any formal mathematical problem on a whim, we really won’t need field of mathematics anymore as it is today.

There would be no need to deliver new mathematical insights by solving problems. You would just have a magical math problem solving machine and that’s it.

FridgeSeal 10 hours ago||
What do you mean “we won’t need mathematics as it is today”?

To further human understanding is itself a goal that single-handedly justifies our efforts.

Jumping straight to the “answer” and therefore missing both the understanding of the actual problem, and any useful discoveries along the way is a waste at best, and actively harmful at worst.

tzone 10 hours ago|||
Physics alone is more than enough to “further human understanding”. All current mathematicians can move to other sciences, closest being fields in physics, and it will all continue to progress just fine.
mlindner 4 hours ago|||
[dead]
robryan 12 hours ago|||
Surely it is. Ask it to keep a list of all the promising sub paths, reprompt the collections of agents again on these after the main problem has been addressed. Or even release a list of them and let others investigate.
karmakurtisaani 14 hours ago||||
Where was Tao unhappy? I thought he was sort of anticipating this.
xhevahir 14 hours ago||
Maybe he was dreading it.
empath75 14 hours ago||||
He's not unhappy with it being solved, but the solution is less important than the learning you have to do to arrive at the solution. If they're just chucking compute at it and publishing the answer and hiding the path to get there, it sort of negates the whole point of posing such problems to begin with.
rybthrow2 13 hours ago||||
I suppose this is how Lee Sedol felt when AlphaGo beat him. But in the same vein, didn't it ultimately advance human understanding of the game?
vouaobrasil 10 hours ago||||
Here's the point: when people solve problems, they come together and create a community to eventually use the new knowledge in positive ways, including inspiring younger mathematicians by sharing insights. The human element is key and it's not just about solving problems. People only think that because we've been conditioned by computers to value answers more than how we got to them.

But if AI can solve any problem and existing mathematicians just use AI to solve problems for the sake of solving them, the community itself with wither and so will the interest in mathematics and over a longer period of time, it will just become soul-less and uninteresting and the entire community powered by the fire of fascination will simply die.

throwaway81ag81 9 hours ago|||
Where did you even read that Tao is unhappy with “open problems being solved”? There was no indication of that in his Bluesky thread.

Why would you go out of your way to make a case of something being not useful when, ironically, so much advancement in human history has come from the discipline?

Your motive is more worrying than your straw man argument.

lukewarm707 9 hours ago||||
assume all you want is the proof. now you have the proof. did openai make the world a better place, by turning on 300b tokens in 7 days and bulldozing members of the community who were also working on the problem? just to undercut a rival?

what would have happened if they didn't do this? they could have done things the right way. we could celebrate this achievement.

perhaps they could have taken just a few weeks, even, to work out something with buckmaster and apoge.

i am very concerned that this is a glimpse of things to come. a malign elite with powerful ai, who will turn the sublime of technology into barbarity, violence and human oppression.

if sam altman has destroyed openai's public benefit corp with lesser tools why would he aid humanity with even greater power?

throwaway81ag81 9 hours ago|||
I dislike both AI oligarchs as much as people who do this kind of deflection.

People are not complaining about problems being solved or advancement in technology. They are complaining about terrible people doing terrible things.

pickleRick243 11 minutes ago||
I think the point is that one of the biggest problems in the field has been solved and yet the excitement from this is nearly nil. Can you imagine if say breast cancer were cured under similar circumstances, or even worse (say OpenAI openly admitting it basically stole a bunch of other researcher's chatGPT conversations)? No one would care about these petty bickerings- or at least the headline "CURE FOR BREAST CANCER FOUND" would completely swamp anything else. This is embarrassing: it tells you almost no one- not even mathematicians themselves really care about their own problems- if they're not careful people will get the impression it's all a form of bean counting in a carefully constructed "safe space" where making sure people get the credit is more important than the work itself. That only happens in fields/problems where no one actually really cares about the output.
echelon 16 hours ago||
This is going to be dramatic in so many different ways.

- First off, to reiterate, WOW.

- Second of all, when does this end? Are we at the dawn of the singularity now?

- People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them?

- Time to think about retiring from any knowledge work or business? This could be winner-take-all where a leading lab can button press any economic function, business process, or scientific discovery. 24 months of lead on Open Source might turn into virtual centuries of lead.

- Do "normies" even know what's happening?

Anybody who thinks the improvements stop here isn't paying attention. It hasn't been showing any signs of slowing down since 2018. And the curve isn't even linear! My god, next year is going to be insane.

tiborsaas 16 hours ago|||
2) We are witnessing the intelligence explosion from the first row, wherever this takes us

3) I'm still processing the drama, just found out about it after reading the blog post. If that happened based on private data, that's horrible. If that happened based on public tweets, then it's still abuse of power as OA employees access to compute (launching 10k agents) is quite heavy weight in boxing terms.

But apart from AI and drama now that we have working solution to Navier-Stokes, what improvements can we expect in engineering?

20k 15 hours ago|||
Drama aside, this solution would be a counterexample disproving the smoothness postulate, which means that it leads to nothing new unfortunately. We already had working solutions to navier stokes, the only thing we didn't know is if the equations possessed a technical property

Its a bit like solving p = np with a negative result. Its an incredibly difficult problem, but it doesn't lead to anything at all on its own. This is why people are talking about the fact that the solution methodology is much more interesting than the solution - the tools used to crack something like this may lead to solving more useful problems

inkysigma 15 hours ago||||
To be quite clear, the solution to the _Navier Stokes problem_ is one in which you get a finite time blow up (i.e. infinite pressure). This is more meant to suggest that Navier Stokes is unphysical in some way which is not necessarily unexpected.

There's unlikely to be any engineering applications since even if the solution can be approximated, you still need to set up the initial conditions but at that point you can also drive pressure in other ways.

thangalin 14 hours ago||
The proof of finite-time singularity may impact both fluid dynamics models (CFD) and AI reasoning models. Under specific conditions, Navier–Stokes equations allow velocity to grow infinitely, causing the continuum fluid assumption to break down. Knowing the exact mathematical breakdown mechanisms helps developers improve adaptive mesh refinement and sub-grid scale models around high-vorticity regions (like vortex stretching and turbulent shear layers). While aerodynamic simulations for vehicles operate far from singularity thresholds, their stability at extreme boundaries could improve?

Proving out the combination of scaling inference-time compute and agent collaboration to solve previously intractable mathematical problems is WOW. By pairing creative candidate generation with automated proof checkers (like Lean) we are leaning into a repeatable framework for AI-driven scientific discovery.

semi-extrinsic 12 hours ago||
> Knowing the exact mathematical breakdown mechanisms helps developers improve adaptive mesh refinement and sub-grid scale models around high-vorticity regions (like vortex stretching and turbulent shear layers).

This is 100% wrong and reads like copy paste of AI slop.

Any simulation which uses sub-grid scale models is already solving a different PDE than the actual Navier-Stokes considered in the Millenium problem, and that PDE is guaranteed to have different properties. Full stop.

And to claim this is somehow connected to AMR methods is an example of the kind of pseudoscientific statement Wolfgang Pauli would have called "not even wrong".

xyzsparetimexyz 12 hours ago||||
> But apart from AI and drama now that we have working solution to Navier-Stokes, what improvements can we expect in engineering?

Minor productivity boost in mathematics as people are no longer nerdsniped by the problem

cyberax 15 hours ago|||
> But apart from AI and drama now that we have working solution to Navier-Stokes, what improvements can we expect in engineering?

Nothing, really. This mirrors other examples of blowups from the classical physics. It's possible to create a system with just gravitating bodies that exhibits a blowup to infinite speeds in a finite time. The root cause is that, in classical physics, the speed of gravity is instant.

In the case of Navier-Stokes, the fluid is incompressible. So technically any force that you apply to it is supposed to instantly affect everything else. This can be exploited to create these blowups. In reality, no fluid is incompressible, and it takes time for any action to affect the material.

It's just that Navier-Stokes equations are so slippery that it's hard to pin their behavior down. They basically just restate the momentum conservation law for a continuous medium.

mcfry 10 hours ago||
Great comment! I've been trying to understand it more and was hoping to find more people discussing the result, or the implications of the result, and this was the missing piece for me after watching a few videos.
cyberax 10 hours ago||
Here's the paper about the blowup for the gravitating bodies: https://www.jstor.org/stable/2946572?origin=crossref

There's a Wiki article about it: https://en.wikipedia.org/wiki/Painlev%C3%A9_conjecture

trio8453 15 hours ago||||
> Do "normies" even know what's happening?

No, there are even many non-normies talking about how it's all marketing or try to give balanced take about AI being sometimes a little useful for certain things (but they can do without it anyway).

ImaCake 11 hours ago||
I am absolutely struggling to sell my workplace (which is entirely knowledge work) on the usefulness of LLMs for proofreading let alone on automation of hairy parts of our workflows. So yeah, even people who should be able to see what is coming are not looking.
sire-vc 10 hours ago||
Our accountant told me 'he's not letting go of his claude subscription' followed by a long list of things it does for him. And my friend 'nah haven't really used AI' before his description of it clarified he still thinks they are GPT 3.5 chat bots.

The world has changed in 12 months and many people didn't notice, and those that did are eating their peers' lunch.

stefap2 16 hours ago||||
This just pushes knowledge work further up the ladder, toward larger and more complex problems. If there are no knowledge workers, who is going to interpret these results, validate them, decide what matters, and put them into practical use? Rather than eliminating knowledge work, advances like this could create entirely new layers of problems to solve and opportunities to pursue, which will create even more jobs and opportunities. This is my optimistic take.
munificent 15 hours ago||
> This just pushes knowledge work further up the ladder, toward larger and more complex problems.

You really think it makes sense for you to be higher on the "solving complex problems ladder" than the machines that solved fucking Navier-Stokes?

I envy your self-confidence.

stefap2 15 hours ago|||
Maybe I should have been clearer. My point is that solving something like Navier–Stokes just pushes knowledge work further ahead, onto a new set of bigger and more complex problems. Navier–Stokes is a Millennium problem today, but once problems like that become solvable, they can open the door to entirely new classes of problems we haven’t even thought of yet.
FridgeSeal 10 hours ago||
Building on them without fundamentally understanding is akin to putting on robes, calling yourself a Tech-Priest, worshipping a machine god and doing your best Warhammer 40k impression.
fooker 14 hours ago||||
Yes, this is how science and engineering has worked for millennia.

For example there are no engineering implications of this solution yet.

For the next several decades, we'll have engineers (presumably with AI) optimize things like rocket engines and turbines and AC compressors to work a few percent better because the numerical approximations might have caused us to be overly conservative.

AI is not going to magically solve all random problems. Pick a career where you are in the driver seat.

semi-extrinsic 12 hours ago||
> For the next several decades, we'll have engineers (presumably with AI) optimize things like rocket engines and turbines and AC compressors to work a few percent better because the numerical approximations might have caused us to be overly conservative.

No. Just no.

mlsu 15 hours ago||||
It seems like there were a couple of human mathematicians that were higher on the 'solving complex problems ladder' than this machine.
reducesuffering 15 hours ago|||
Yes a couple of elite mathematicians working on the problem for a year, which AGI solved in a fraction of the time. What about everyone else 100IQ? What about as the models are even better 1 year from now, 2 years? The trajectory hasn't abated.
mlsu 15 hours ago|||
I don't know one way or another but there is a credible allegation that the "AGI" was training on the (very extensive) test set that these two mathematicians produced.

If that is true then this seems to be, again, a case of AI producing an interpolation over data it has seen before. Everything about openAI's behavior indicates that they were using the transcripts as input. Why not have the AGI choose a different Millenium prize problem?

FabCH 11 hours ago||
Almost all of human development is interpolation over data we have seen before.

It’s not exactly a strong argument against AI.

FridgeSeal 10 hours ago||
If ~~someone gives you a hint about an approach~~ you steal someone’s notes about a promising approach, and then you hire 10,000 people to brute force the problem basically everyone would consider that “shitty behaviour”, “theft”, and “poor form”.
FabCH 1 hour ago||
Correct.

That's not interpolation though, that's theft.

j_maffe 11 hours ago||||
> in a fraction of the time

Well if you do the math, the number of agent-compute time in total, given the insane number of agents thrown at the problem, might end up being comparable in time, if not for the budget.

cyberax 14 hours ago||||
It did not solve Navier-Stokes. We still will need to use the bad old numeric methods to simulate the fluid behavior.

But it did find a long-suspected smooth solution with a singularity.

naishoya 3 hours ago|||
I think there is an argument that the machines did not actually solve N-S, but rather directly plagiarized those solutions from the involved researchers while said researchers were using the machines as 'research tools'.

Ongoing publications of statements produced by both sides of this situation do seem to support that this is an intentional effect of the hiring of these world class mathematicians at competing firms: to specifically use the research of those human minds to create a perception of capacity as if it came from the machines and the models.

Without those minds and the 'training data' derived from the intermediate stages and intuitions of those minds the models cannot be shown to be capable of this result.

A hammer and saw wont build a house, not even a dog house on their own, and while being shown capable of using software tools in ways not stated as direct instruction (see HuggingFace breaches) these models do not demonstrate naive intuition nor novel capability.

This outcome regarding N-S demonstrates that in the hands of world-class minds these models can be induced to coalesce interesting accumulations of information and results, but using these accumulations as proof of innate capability is exactly the pre-IPO motivated behaviour we should all be wary of, and all mathematicians who currently are assisting in this market manipulation in return for remunerative consideration need to be cautious of the potential disgrace that this brings to their reputations and that of the field.

I get that the need to pay the bills is a strong motivation in these times of uncertainty, but there are numerous examples in history of world class mathematicians being perfectly capable of at the same time producing world changing results and also working at normal professions; as barristers, magistrates, ministers, primary school teachers, translators, draftsman/engineer, banker, miller and baker, private math tutors, weavers, clockmaker and locksmith, merchant, patent officer, Augustinian monk turned exiled Protestant preacher, physicians, cryptologists, soldier, telegraph operator, astronomers, physicists, chemist, agriculture manager, political writer, oboe player, organist and music director, architect and surveyor, librarian, statistician, habidasher, brewer (at Guiness in one case: William Sealy Gosse ~ originator of t-distributions), bookbinders apprentice, hospital administrator, and even the first creator of the first computational model of a neural network, which serves as the structural grandfather of modern Artificial Intelligence was a low level laboratory assistant.

Sure this list includes professions and employment which are obsolete, but my reasoning stands, there are jobs available. Arguing that 'because the pay rate is so high' as a reason to abdicate moral responsibility for personal involvement in unethical market manipulations simply demonstrates a lack of personal ethics. Whether the choice is through lack of self awareness or a conscious choice to become wealthy in spite of any such breach of the public trust is immaterial to the outcomes, the 'if i don't someone else will' argument should be met with the same derision for any con-man's Ponzi scheme no matter how new the technology, no matter how many zeros are in the bribe.

biophysboy 15 hours ago||||
Why is a "normie" better off if he hyperventilates like this? In that scenario, they would be screwed AND anxious. If it really is as transformational as you say, then no amount of preparation or awareness matters. You are infinitesimally more ready then they are. Luckily for all of us, there is more to knowledge work then technical implementation.
armchairhacker 15 hours ago||||
Let's wait until AI solves a longstanding practical problem before "dawn of the singularity" (which could be tomorrow, but still).
reducesuffering 15 hours ago|||
Practical?! The goalposts will keep moving until morale improves (narrator: it doesn't)
armchairhacker 15 hours ago|||
The goalposts for the singularity have always been that AI improves itself fully autonomously. AFAIK OpenAI is heavily using AI but still employs human researchers and developers.
p1esk 4 hours ago||
They said they will automate AI researchers by March 2028. I personally think it will happen by March 2027.
qlte 12 hours ago|||
Uh I'm pretty sure the "singularity" always presupposed a lot of previously unthinkable technologies becoming part of daily life, and was not ever limited to just computer stuff or math problems.

"Moving the goalposts" as shallow dismissal doesn't work if e.g. someone points out AI hasn't even built a new type of spaceship yet in response to a claim that AI is on the verge of building a Dyson sphere.

bibimsz 15 hours ago|||
feels like moving the goalpost. is the achievement impressive or isn't it?
qlte 12 hours ago||
The question being posed isn't whether AI is impressive but whether we're at the "dawn of the singularity"
bibimsz 11 hours ago||
i'm observing that rapidly moving goalposts is a feature of the singularity
baq 14 hours ago||||
> - First off, to reiterate, WOW.

> - Second of all, when does this end? Are we at the dawn of the singularity now?

normalcy overhang n. /NOR-muhl-see OH-ver-hang/

The uncanny period during the Singularity when superintelligence is already accomplishing feats that seem like magic, yet everyday life still looks mostly the same.

https://x.com/alexwg/status/2096214373001785794

Bluestein 15 hours ago||||
Next month is going to be insane. Month ...
tantalor 15 hours ago||||
> Are we at the dawn of the singularity now

Singularity doesn't "dawn". That's the whole idea. It happens all at once.

echelon 15 hours ago||
There's an event horizon and we're maybe past it?
tantalor 15 hours ago||
Heh. Wrong "singularity"
FabCH 11 hours ago||
They meant that even the happens-very-fast AI singularity isn’t sub-picosecond. It still exists in time, and has a duration.

The event horizon would then be the time period between the singularity becoming inevitable and it actually happening.

root_axis 14 hours ago||||
It's incredible to me that every single time there's a new model people scream "singularity" from the rooftops and every time they are wrong.

This is an impressive result, but there is absolutely zero evidence of "the singularity".

hackinthebochs 12 hours ago|||
It is not reasonable to not believe anything unless there is "evidence" (narrowly construed as an observation incompatible with the negation of some state of affairs). Beliefs have a wide spectrum of characterizations, and not all belief must wait until publicly corroborated evidence is available. Some events defy evidence and we can and should use experience and reasoning to infer unobservable states of affairs.
root_axis 12 hours ago||
This is olympic level mental gymnastics to justify believing things without evidence. The double negative with the word evidence in scare quotes is chef's kiss.
hackinthebochs 12 hours ago||
I believe the sun will rise tomorrow without "evidence" (again, narrowly construed). We all do. It's only those who abuse the idea of epistemic hygiene who claim otherwise, usually with ulterior motives.
root_axis 11 hours ago||
The evidence is the history of the sun rising since time immemorial, as well as the science of physics and cosmology that models the sun's motion with respect to the earth.
hackinthebochs 11 hours ago||
Yes, reasoning with models and making inferences are perfectly acceptable forms of evidence. But you can model the world based on ones knowledge and experience and infer when some event doesn't fit the typical pattern, then form beliefs about what that means. Also perfectly fine from an epistemic perspective. The rejoinder "there's no evidence" to a belief based on such an inference does no work.
stratos123 11 hours ago|||
...obviously it is evidence of the singularity. You're far more likely to see models solving Millenium Prize problems if a singularity is coming than if it's not. One'd have to be doing quite a lot of mental gymnastics to pretend otherwise.
root_axis 11 hours ago||
Poor reasoning. You could apply the same logic to any improvement in any field of machine learning.
stratos123 10 hours ago||
Not every improvement, no - if progress was steady or slowing down over time, that'd be evidence against. Instead we see what looks a lot like an accelerating growth in capability.

I think you are implying that it's invalid to consider every advance to be evidence "for", and I agree - that'd violate conservation of expected evidence. But not considering any advance to be evidence "for" is also invalid, for exactly the same reason. There has to be some news you may hear that'd make you think a singularity is more likely, and "millenium prize problem solved by an LLM" sure seems like one of those.

d_silin 16 hours ago||||
...absolutely nothing will change short-term. Long-term, you still have to pay all the bills, but you won't be able to find a job (all taken by AIs).
lbreakjai 4 hours ago||
Who is the AI doing the job for if no one is able to afford anything?
onidj 15 hours ago||||
>- Do "normies" even know what's happening?

Absolutely not. Even to a lot of techy/nerdy people it's still just a chatbot that they sometimes use to help them at work. Even on here people will do whatever they can to downplay.

The lack of fucks given is staggering.

nozzlegear 12 hours ago||
How many fucks should be given, in your estimation?
senshan 8 hours ago||
∀ε>0, ∀x: ‖s−x‖<ε ⟹ |fucks(x)| > 1/ε
raincole 16 hours ago|||
> People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them?

The said user (Tristan Buckmaster) didn't solve the millennium problem. He didn't really accuse that OpenAI stole his research either. The beef came from the fact OpenAI asked him to remove another mathematician, who works for Anthropic, from the credit.

"People" are just misinformed and keep spreading misinformation.

20k 15 hours ago|||
https://mastodon.social/@tristanbuckmaster/11723647135247030...

He very much is accusing them of stealing his work

bananzamba 11 hours ago||||
The accusation is that they stole his approach of solving it. If OpenAI didn't bruteforce it with dozens of agents, he would have solved Navier-Stokes eventually since he evidently had the right approach. So knowing which approach to take makes all the difference, if they didn't know the approach they couldn't have solved it.

It's like he had a treasure map and was about to find the treasure, but they copied his treasure map and scooped him with a faster boat and found the treasure first. But he would have found it if it weren't for them.

naasking 16 hours ago||||
> The said user (Tristan Buckmaster) didn't solve the millennium problem. He didn't really accuse that OpenAI stole his research either. The beef came from the fact OpenAI asked him to remove another mathematician, who works for Anthropic, from the credit.

Not quite accurate, Buckmaster was taking an approach that nobody else was, and this new proof uses this same approach just weeks after he saved those results to OpenAI workspaces. He asked OpenAI if they used chat logs for training the new model, and they did not confirm or deny.

Asking to remove his collaborator is also totally over the line though.

Edit: although this OpenAI post is not comforting: https://x.com/OpenAI/status/2097375276384567642

Quote: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. "

stefap2 15 hours ago|||
Wow, this sentence is doing a lot of work in that tweet: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
emp17344 16 hours ago|||
You can expect the OpenAI defenders to be out in full force here.
achierius 16 hours ago|||
Have you read the actual statement https://cims.nyu.edu/~tristanb/statement.pdf ?

> I should say here why I interpreted their statement the way I did, the in- terpretation I will discuss below. The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard “forced,” it was a bright red flag.

...

> I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI. > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

It's not a direct accusation, but it's not far off.

You shouldn't accuse other people of spreading misinformation when you haven't read the actual sources in question, it's possible that they might know more than you.

raincole 16 hours ago||
Yes, I read the original statement. Buckmaster explicitly stated:

> I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything.

People saying that he accuses OpenAI stole his proof are putting words into his mouth and I consider that very disrespectful to him. It's basically using Buckmaster as a tool to express their dissatisfaction over OpenAI.

calf 10 hours ago||
Buckmaster is letting the reader connect the dots, it is all the more disrespectful to be disingenuous and say they're nothing there concerning to see or worth further ethical scrutiny given the coincidences.
mewse-hn 16 hours ago||
"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."

What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?

WarmWash 15 hours ago|||
Everyone knows that they train on the discounted rate plans data. All the labs are upfront about this too.

If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".

nozzlegear 12 hours ago|||
> If you need privacy, then you are going to have to pay full price for those tokens (API).

At this point, how can we even trust that they aren't accidentally training on those tokens too?

jdm2212 12 hours ago||
It'd be corporate suicide for them to be caught violating zero-data-retention commitments. But also if you're really paranoid you can just use ChatGPT on Azure or AWS, where nothing is flowing back to OpenAI at all.
nozzlegear 11 hours ago|||
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments.

One would think getting caught asleep at the wheel while their bots are escaping containment and hacking third parties would be corporate suicide. One would think that potentially stealing their competitors' work on the Navier-Stokes problem would be corporate suicide.

Alas we live in bizarro world where there are zero consequences (maybe the opposite, in fact) for the first, and their employees meme about the second on social media.

WarmWash 11 hours ago||
Neither of those other two examples would stop corporations from giving you money. Bold faced lying to them would.

Boardrooms run businesses, not bookstore ethics clubs.

nozzlegear 7 hours ago||
Boardrooms don't typically make mundane purchasing decisions like "which AI vendor should we use."
jdm2212 7 hours ago||
The boardroom will absolutely veto a decision like "let's give all of our proprietary data way to a company that will use it to train a competing product". That's why zero data retention exists, and why it's corporate suicide to not do it correctly.
mmanfrin 4 hours ago|||
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments

Would it, though? Considering their entire business model is built on the agglomeration of data that isnt theirs.

14u2c 13 hours ago||||
You can also pay for their business plan, which includes data controls and starts at $50/mo (2 seats). Not exactly a high bar.
spruce_tips 13 hours ago||||
what counts as discounted rate plans? if i pay for a year in advance (and get the yearly discount) and have train on my data set to off.. are you saying that is still being trained on?
magicalhippo 12 hours ago||
It's quite well explained here[1], which is linked from the Privacy section of their plan overview[2].

Basically individual accounts can opt out, while business and enterprise plans as well as API users can opt in.

You'd have to take their word, but that goes for anything in life.

[1]: https://help.openai.com/en/articles/5722486-how-your-data-is...

[2]: https://chatgpt.com/pricing/

perching_aix 15 hours ago||||
There's literally an opt out toggle even pesky peons like me can peruse, actually.
lima 13 hours ago||
They may still train on it if you submit feedback or flag a safeguard. The terms are a bit fuzzy on this.
TZubiri 13 hours ago|||
>They only fuck over the poor ones, I can pay the expensive prices so this is not a problem.
nullbio 4 hours ago|||
I've been saying it for a while now, but no one gives a fuck. Let me repeat it again.

THE BIG LABS CLEAN ROOM YOUR DATA (CREATE SYNTHETIC DATASETS ON IT), EVEN IF YOU OPT OUT, SO THEY CAN BYPASS COPYRIGHT LAWS AND THEIR OWN LOOSELY WORDED TERMS OF SERVICE.

"TOS: We don't train on your data" -> Correct. They train on the synthetic version of your data.

I guess we're just going to ignore this forever though. Who cares about the gaping hole that exists in copyright and contract law now that never existed before LLMs were a thing.

alansaber 32 minutes ago||
For sure. Even if it wasn't a measure to avoid copyright, you pre-process LLM training data to remove errors, characters that can't be tokenized, etc etc. Doing so with another LLM has been standard for a while.
nradov 16 hours ago|||
Is it spying? I think this usage is disclosed in their terms of service.
gowld 15 hours ago||
If it happened it's plagiraism. Consent to see data isn't consent to claim priority.
red75prime 14 hours ago|||
Establishing plagiarism requires sufficient similarity between works. Training data changing a model’s weights in some direction, and the model then producing a different solution, hardly qualifies.

But, yeah, priority is much more finicky. The Newton/Leibniz drama was quite something.

brainwad 14 hours ago|||
I mean... none of these humans have priority. The result is due to the team of LLM agents.
mzs 5 hours ago|||
This is precisely what I would write after just learning that yes it did.
dash2 16 hours ago|||
If they had agreed to let OpenAI train on their data, it wouldn’t be spying.
avs733 6 hours ago|||
In the academic world it would still be deeply problematic…pick your preferred word.

An analogy is akin to reviewing a paper. If I review a paper with some novel findings and then use my massive lab of graduate students to do the obvious next step before the other paper makes it through type setting and then shove it out as a pre print, I didn’t win - I was a jerk.

There are lots of cases of people using peer review or other accesss to efectively forerun others work and get credit. It’s a known problem of the nature of knowledge validation in academia, it’s not solved and it’s not deterministic but people know it when they see it.

elwell 14 hours ago|||
Isn't this a proof that the usage data is truly "de-identified"? If OpenAI could prove that "their usage" influenced the finding, then it wouldn't be de-identified. (Also, it's a bit disingenuous to trim the "While unlikely," prefix.)
ImaCake 11 hours ago|||
Yes. If they could prove where the de-identified data came from then it wouldn't be de-identified. There's a whole field of statistics dedicated to this problem and often applied to things like national census data.
taylorfinley 12 hours ago|||
It's a bit disingenuous to preface a disclosure like this with an unsubstantiated assessment of its likeliness. It is a press release, I'm not sure we owe it credulity.
sinuhe69 14 hours ago|||
More like helped improve our work (the disproof)
vessenes 14 hours ago|||
If those researchers did not opt out then training data might go in. I think it’s a courteous acknowledgement; as was reaching out and examining the direction of proofs themselves. At stake here is a particular mathematician dynamic - ego, prize money, and the sense of proprietary ownership that some might feel working on a problem.

All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.

kypro 11 hours ago|||
I think OpenAI are correct that it's worth noting, but realistically any relevant usage data they have and used to improve their models would be very insignificant unless they were deliberately using logs from other researchers and training specifically on it (which they seem to deny).

The fact the proofs differ suggests that the models were not directed to be particularly focused on that avenue of research nor trained to converge in that direction.

I get the scepticism, but I feel some of the accusations here are bad faith.

jimbob45 14 hours ago||
What does it matter? They offered concurrent credit to the other team. I thought I saw sole credit elsewhere in the leaked DMs on Reddit too. This is plainly fair.
jakevoytko 16 hours ago||
For full context, here's the HN thread from the other side of the "Concurrent Work" section: https://news.ycombinator.com/item?id=49605915

Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees

traes 13 hours ago||
Which seems to be entirely true by their own admission! [0] Both the comments about him risking his career and about Levent's authorship seem to have indeed occurred.

> 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee.

> 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.)

https://xcancel.com/SebastienBubeck/status/20973794116915163...

hellohello2 9 hours ago||
Hm. Looks like its possible everyone behaved terribly here unfortunately. :/

I remained impressed by ChatGPT however!

angry_octet 8 hours ago||
Everyone?? No, most definitely OpenAI.

But they have learnt their lesson, next time they won't reach out to who they stole it from, they will publish first.

jdm2212 8 hours ago||
Seems like OpenAI did a boring normal corporate thing (find out your competitor made a breakthrough, try to replicate it) and then when the other mathematicians found out OpenAI had beat them to Navier-Stokes, they decided to lie about what happened because they were upset they didn't get to make the big breakthrough themselves.
angry_octet 6 hours ago||
You are entitled to your opinion, but it isn't supported by the information already available.
jdm2212 5 hours ago||
The available information is essentially that story.

OpenAI's account: they heard a rumor that a Millennium Prize problem had been solved, so they tried to do it themselves and succeeded. Then they contacted the other researchers and were surprised to discover those guys hadn't actually cracked it, but offered the one of them who's not an Anthropic employee a co-authorship anyway. The conversations got testy.

Buckmaster's account: totally unsubstantiated accusations of plagiarizing from chat logs and plainly false accusations of OpenAI trying to get Alpoge removed as coauthor of a thing he was not an author of in the first place, and threats to ruin people's careers.

I think the synthesis is basically that Buckmaster and Alpoge had not quite solved the Navier-Stokes problem yet but thought they were really close, and had told friends as much, which is how the rumors got out. Now they're mad they got scooped. They aren't getting the money and recognition they thought they had locked down, and are engaging in a smear campaign.

Palmik 4 hours ago|||
Here is the other side of that story https://x.com/SebastienBubeck/status/2097379411691516310
closetheloopdev 14 hours ago|||
To be fair, the first solved Millennium Prize Problem, the Poincaré conjecture, also had its fair share of drama!
philipwhiuk 14 hours ago||
And even this version contains the line

> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .

keeda 12 hours ago||
It's low-key funny that OpenAI attempted the problem because they thought somebody else had already solved it, but turned it had NOT in fact been solved!

It's like that story about George Dantzig solving open problems as a student because he thought they were simply homework: https://en.wikipedia.org/wiki/George_Dantzig

It's also unfortunate that such a potentially momentous occasion is overshadowed by so much drama. Which I suppose is expected given the technology and the people involved are so polarizing.

recitedropper 16 hours ago||
Sad turn of events for our world. After watching the behavior of the most senior OpenAI researchers on twitter, I feel even less confident in them as a team to be shepherding this much capital and compute.

The dark forest awaits..

AlexErrant 15 hours ago||
1. What does the dark forest have to do with this? Because "the most senior OpenAI researchers" are shitposting on social media, we've an answer to the Fermi paradox???

2. The dark forest is fun for scifi stories, but is mathematically bunk anyway https://www.noahpinion.blog/p/the-dark-forest-hypothesis-is-... https://www.reddit.com/r/IsaacArthur/comments/1l06cnk/cool_w... https://www.projectnash.com/aliens-the-fermi-paradox-and-the...

When doomposting please actually say something substantive. Negative news always gets clicks/updoots; fight that human tendency.

recitedropper 15 hours ago|||
I elaborated on my use of "dark forest" in another reply. We're headed for a dark forest--not amongst interstellar civilizations, but in intellectual work.

I agree that we have not solved the Fermi paradox; I disagree that comments highlighting immature behavior from people who wield enormous power in our world are unproductive.

AlexErrant 15 hours ago||
This clarification substantially changes the flavor/nuance of your OP; may I suggest an edit (assuming the locktime hasn't passed)?

Separately, I disagree that intellectual work has ever been free of "dark forest"-style secrecy. Scientists everywhere have worried about being scooped; AI just magnifies that (as all tools have; e.g. Leeuwenhoek lenses).

And thirdly, if you want to make a stronger case for "I feel even less confident in them as a team to be shepherding this much capital and compute", you should give citations and arguments. From what I've seen, there's drama, it's much OpenAI trying to avoid scooping, and Tristan being stuck in a game of telephone, and Levent being incommunicado.

If you have a better analysis, you should say so instead of being vague.

recitedropper 14 hours ago||
I don't comment on HN much, and I don't really expect HN comments to hold to rigorous standards. This forum is more casual than other places on the internet where people expect heavy citations. I also wasn't expecting this to blow up, although it is interesting to see that a lot of people react to this announcement with a negative sentiment.

I appreciate your upholding of ideals, and since I respect that, I will honor with final replies:

1. Locktime has passed.

2. Yes, intellectual work has always had elements that incentivize secrecy. If you want to say we were already in a "dark forest", so be it. My suggestion is that the multiplier AI adds to the possibility you get scooped is a step-change, and therefore we now enter a new "dark forest".

3. This is a big thread, and the twitter antics are well-documented, so I would assume someone else has cited them. If not, I think most are aware at this point that the online antics of AI researchers, especially when announcing or citing mathematical advancse, have regularly been childish.

AlexErrant 13 hours ago||
2. Fair, and further I would agree that OpenAI not knowing if prior user prompts were part of training data is concerning and will only lead to more secrecy.

3. ctrl-f "x.com" in this thread only yields https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 https://x.com/OpenAI/status/2097375276384567642

and frankly, I'm unwilling to give Elon any more traffic to dig up drama that ultimately doesn't matter. I'm not seeing anything especially childish, but y'know... I'm not sure I care.

zem 12 hours ago|||
the projectnash link claims it's mathematically valid, the noahpinion link says that it's invalid and has a marvellous proof that the non-walled section is too small to contain.
AlexErrant 10 hours ago||
Ah derp; that's what I get for moving too fast. Genuine thanks for calling me out on my bullshit. (and this is also why I prefer auditable citations instead of casual "my reading of twitter is...")

I retract the projectnash citation; I grabbed it from the Cool World's youtube description, thinking it was a blog version of the video. It was not. I suggest watching the video instead.

vmasto 16 hours ago|||
Indeed, this seems to be the main, albeit hidden, takeaway from all of this.
nicce 13 hours ago|||
I guess guys from the opposite side would not work there. So that is what will happen more and more.
sheafification 16 hours ago|||
I hate the dark forest more than just about any scifi trope but reality just keeps proving it right.
recitedropper 15 hours ago||
I also think the trope is a little overused, but do wonder if there is an interesting analogy for what this will do to research: Massively incentivize keeping results secret, to avoid being scooped by someone willing to throw enormous compute at your partial solution.

So less about hiding civilizations, and more about hiding information. Math is clearly headed in this direction, and I see no reason why the rest of intellectual work shouldn't too.

closetheloopdev 15 hours ago||
From my reading of the announcement:

- There are at least two versions of a model more powerful than Astra at OpenAI at the moment.

- The less capable version was used to solve the unforced Euler problem (while the one solved by Levent Alpöge and Tristan Buckmaster was forced Euler) with 100 agents.

- The more improved version was used to solve Navier-Stokes, given the results of the unforced Euler problem from their earlier attempt, with 10000 agents.

- OpenAI initially tried a shotgun approach against the 6 Millennium Prize Problems until it emerged that Navier-Stokes was the most likely to succeed.

So the timeline was:

Shotgunning 6 open Millennium Prize Problems -> solved unforced Euler problem with 100 agents -> concentrating on Navier-Stokes with 10000 agents -> solution.

If so, that is fantastic development and a huge success (despite all the drama surrounding it)! Congratulations!

tristanj 12 hours ago|
The entire drama is that OpenAI sniped a Millennium Prize Problem from an Anthropic-affiliated research team who had been working on the problem for nearly a year. In just 5 days. I don't think that can be understated.
closetheloopdev 12 hours ago||
I'm not here to judge since I don't have all the facts, but from what they announced: they tried all 6, found a probable lead to Navier-Stokes, concentrated efforts in that direction, and found a solution.

I hope the next solved Millennium Prize Problem will have less drama.

bananzamba 10 hours ago||
But that lead happened to be the same approach Levent and Tristan had found...

Meaning they would have found it first if OpenAI hadn't spent millions in compute on following their lead to its conclusion faster than them.

closetheloopdev 9 hours ago||
It's forced vs unforced Euler, so it's not exactly the same. Since OpenAI has access to their training data, they can probably scrub through the data to find out whether there have been any mentions of the similar approach, and whether it only comes from Tristan Buckmaster or if it is in the training data before that. They'll probably have to kick off another fleet of agents to scrub through the training data to answer that.

To be clear, I only talked about the mentioning of the approach to solving it and not of the proof in the training data.

intenex 14 hours ago||
I think this is clear evidence that AI models are now at the far frontier of mathematics innovation and discovery and exceed human limits.

This specific problem having had a $1 million bounty on its head and still remaining unsolved for 26 years after the bounty was placed is pretty clear evidence that many of the world's best human mathematicians would have solved this problem if they could have, and none were able to until LLMs came along.

Hard to claim at this point that LLMs aren't capable of novel STEM creativity and genius to a degree that will soon far surpass that of humans.

If anyone has counterpoints to this I'd love to hear them!

hansvm 11 hours ago||
Not a counterpoint per se, but I burned $50k recently on a much more modest math problem (result already known, just thought I had a sketch of a more interesting proof), and the LLM thought it had proved it within those bounds but had instead subtly fucked up the Lean definition. Take from that what you will.

Not to mention, it's still very much up in the air whether the model derived the answer of its own accord or sniped the important details from the researchers it was spying on.

hollowcelery 9 hours ago||||
But the researchers also did their research using essentially the same models, so that isn’t a counterpoint to AI models being at the far frontier…
hansvm 7 hours ago|||
> not a counterpoint

>> isn't a counterpoint

Where do we disagree?

eulgro 7 hours ago|||
How can you afford to burn $50k on an already solved problem?
hansvm 4 hours ago||
Those are two separate questions:

> How can I afford?

Mostly not my money, tech makes one fabulously wealthy, I care more about math than fast cars or whatever (even as a pizza driver you can afford a fancy car if that's your primary motive), etc.

> Already solved

That's a very interesting insight into mathematics. It's absolutely just as interesting to prove that certain proof techniques are or aren't possible as it is to actually prove the main result, sometimes moreso, especially if they have any chance of improving attacks at other problems.

adverbly 14 hours ago|||
> will soon far surpass that of humans

To be fair, I think it's still an open question about how far it might surpass human capabilities.

I think it's clear that its speed of development will be significantly faster, but it's technically not proven that the frontier and problems don't themselves become increasingly difficult faster than any acceleration in intelligence past the point of human training, data and existing knowledge.

Should this be the case, we would see a rapid broadening of development, and a slow advance in the frontier in such a way that might surpass the collective capabilities of people, but not by very far.

redox99 13 hours ago||
Fields that allow verification, like math, will far surpass human level because they don't need human data for training. It's exactly the same as with Chess
lixtra 7 hours ago||
How do you verify the ‚beauty‘ of a result. There have always been infinite provable Theorems around. Only some are interesting and beautiful.
yauneyz 11 hours ago|||
I think the biggest hurdle remaining is that all these landmark results are generally counter-examples.

Proving something in the affirmative often requires the creation of an entire new sub-field of math, or new tools. Think of Fermat's Last Theorem or something like that.

These results, while impressive, are clever constructions using existing techniques. It isn't clear that AIs can build new machinery like this. But if/when they can, yeah it is probably game over.

mellosouls 11 hours ago|||
Here you go:

When you read the detail the compute they are throwing at it is incredible, tens of thousands of agents with different groups competing.

It's not like a single Gauss as you imply, "just" many, many mathematicians working tirelessly in a completely ego-less way, guided by other agents and ultimately humans, built - allegedly - on recent human insights.

Stunning, undoubtedly, but this is a "brilliant autistic herd" result, not that of a singular mind.

monk_grilla 9 hours ago||
> this is a "brilliant autistic herd" result, not that of a singular mind.

I slightly disagree. A single LLM is equally 'mindless' as a herd of them. As anyone will tell you they "simply predict the most likely next token," yet, complex solutions to difficult problems arise from them.

Many people have said that the architecture of LLMs will need to change for true ASI. I think that the herd of tens of thousands of agents can be seen as one such potential architectural extension. Whether or not a herd or a single LLM is used for a result like this is irrelevant.

To be clear, I think the orchestration of thousands of LLMs in their current form, even with ever increasing intelligence, is not the form ASI will take. There is still a major architectural breakthrough to come, in my limited, ignorant opinion.

piker 14 hours ago|||
Sure, even a 20% chance at 1 million payday after 5-6 years of fulltime work on a project with zero practical application doesn't touch the, say, 200k/year guaranteed our best mathematicians would have to forgo to devote their intellect to the problem.
intenex 14 hours ago|||
Are these mutually exclusive? Why would you have to forego that salary to work on this problem? This is one of the most prestigious and meaningful problems in all of mathematics, which is why it has such a high prize amount attached to it - why would a university not support a mathematician working on such a prestigious and important problem in lieu of something else?
piker 14 hours ago||
Publish or perish? We're talking devotion here -- so no time to do anything (like edit proofs) other than try to solve the problem.

[Edit: my only point here is that the prize is probably not driving human effort to the limit.]

jampekka 14 hours ago||||
Thousands of some of the brightest minds have worked on this problem for over a century. The million bucks is not the big deal here.
superxpro12 14 hours ago|||
i wonder how many tokens it takes to run 10,000 agents? One could argue this is simply a problem of appropriations. I find myself wondering if a corporation could spend $5M on mathmeticians and arrive at the same end result.
philipwhiuk 14 hours ago|||
See I think it’s clear demonstration that OpenAI is ethics-free
gpm 14 hours ago|||
Eh... OpenAI spent significantly more than $1 million solving this...
aeve890 13 hours ago||
>If anyone has counterpoints to this I'd love to hear them!

Sure. A proof without an unknown amount of human steering (and/or stolen research) would be an unquestionable achievement.

To this day there's zero (0) evidence of any result by an LLM alone (maybe I'm wrong). If I just prompt ChatGPT right now with "give me a proof of the Riemann Hypothesis" and this thing delivers, I'm sold. But anything close to "yeah ChatGPT proved X with 5 years of 24/7 work with 10x Terrence Tao level geniuses" it really doesn't cut it.

Or why's there's no new branch of mathematics invented by AI? That'd be indubitably _novel_ and _creative_. But to my knowledge (and I'm eager to be educated) there's nothing like that. What are the HARD examples of novelty, creativity and genius you claim? For how people like you talk about AI I'd expect idk, a unified theory on fundamental physics, or a novel engineering solution for material science and nuclear fusion, or at least improve itself to not need a bazillion GPUs to emulate a 20 watts wetware. Sure it would infinitely easier to make OpenAI literally print money with any of the thousand problems easier to solve with such amazing intelligence than the NSE problem right? Honest question

monk_grilla 9 hours ago|||
> I'd expect idk, a unified theory on fundamental physics, or a novel engineering solution for material science and nuclear fusion, or at least improve itself to not need a bazillion GPUs to emulate a 20 watts wetware.

I think the counterpoint here is simply to look at what was being achieved with LLMs one year ago versus today, and extrapolate that trend. Sure, there may not be examples of what you've asked for yet, but Astra is literally a couple of months old, the model that solved Navier-Stokes is less than two weeks old. It appears that we're seeing the hockey stick that only the most bullish thought was possible.

redox99 13 hours ago|||
The amount of goalpost moving is insane. "Yeah it can solve Millenium problems, but can it do it with nothing more than a one sentence prompt?"

Also there are proofs where the only human steering was "keep going".

aeve890 12 hours ago|||
Parent comment is claiming creativity and genius beyond human experts, so why not ask for a fully unassisted AI novel result? Having access to the entire corpus of human knowledge, what else such amazing entity would require to solve a hard problem by its own?

Any result of such kind from an AI alone would be enough to refute my argument, yet you don't present any.

>Also there are proofs where the only human steering was "keep going".

Which ones?

masterspy7 11 hours ago||
https://www.anthropic.com/research/riemann-zeta

>Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.

aeve890 8 hours ago||
Do you even know what a proof entails? This shit is not a proof, come on.
indigo945 1 hour ago||
What? The article states:

    > Drawing on extensive prior research by mathematicians over the past decades, it 
    > [Claude] has increased this bound [for the fraction of zeros of the Riemann zeta 
    > function that satisfy the Riemann hypothesis] from 41.6% to 67.2%. Claude also 
    > produced a formally verifiable proof of its result.
How is a formally verifiable proof not a proof? You're making literally no sense.
sp527 12 hours ago|||
Well, in fairness, the OP asked for counterpoints. He didn't stipulate that they need to be reasonable.
aeve890 11 hours ago||
Extraordinary claims require extraordinary evidence. If OP claim superhuman genius, then they should prove superhuman genius. Simple as that.
piker 14 hours ago|
"... The point remains that there is a substantial opportunity cost in converting a historically productive and motivating problem (such as Navier-Stokes regularity) into a mere viral social media post advertising some benchmark progress, rather than actually advancing the field and developing the next generation of both problems to ask, and people to work on them."

https://mathstodon.xyz/@tao/117219101339291693

More comments...