Top
Best
New

Posted by milkshakes 6 hours ago

Ten advances in mathematics and theoretical computer science(openai.com)
346 points | 630 commentspage 3
randomizedalgs 1 day ago|
After skimming some of the writeups, I'm surprised that the frontier internal model still writes just as poorly as Sol.

Maybe good AI paper writing is further away than I thought...

QwenGlazer9000 6 hours ago|
You mean we're still gonna be employed doing the boring part while AI gets to do the fun part?

I'd honestly rather they just automate every job at that point.

readthenotes1 2 days ago||
I wonder if Erdos would be saying " It's fine that y'all are answering my questions, but who is asking better questions??"
macleginn 2 days ago||
I am duly impressed by the powerl of the nameless internal AI, but not a single human contributor's name listed anywhere? Did someone at least make this model a coffee?
drdrey 2 days ago||
> The results were achieved by an internal version of Astra, our next major model.
zogomoox 2 days ago||
surely some human regularly typed "think deeper, make no mistakes".
zardo 2 hours ago||
My grandmother is very sick and the doctors need this proof to help her.
danielrmay 2 days ago||
I'm enjoying learning about these hard problems, but this line about credit made me chuckle:

> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness

Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?

rencrisa 2 hours ago||
It seems that a lot of folks misunderstand the guarantees that lean provides.

I just want to state that having "lean proofs" that build (checks) does not mean the actual real theorems we care about hold. Ignoring lean kernel bugs, ultimately a human (not an agent) has to verify the lean encoded theorem statements (specs/specifications), that the lean proofs are checked against, indeed correctly encode the real theorems. For non-trivial theorems such as these, this is an arduous and tricky task where even a little mistake could be fatal. AI generated lean encoded theorems can be huge and difficult to understand. I wonder if anyone reputable has audited these specifications.

DroneBetter 2 days ago|||
well, a bug in the Lean kernel was discovered last week by way of an LLM tricking itself and its handler into believing it had found a non-constructive proof of the existence of a nontrivial Collatz cycle, see https://infosec.exchange/@0xabad1dea/117002106099986943 and https://lipn.info/@mevenlennonbertrand/116997917683191056
traes 2 days ago|||
That seems to have been more of a sensationalized joke. Even your link has a disclaimer in it now. Read this chat from the researcher who did this:

https://leanprover.zulipchat.com/#narrow/channel/270676-lean...

jibal 2 days ago||
It's not at all a joke ... that's a severe misunderstanding of the context.
traes 2 days ago||
There is no evidence that I can find for the claim "a bug in the Lean kernel was discovered last week by way of an LLM tricking itself and its handler into believing it had found a non-constructive proof of the existence of a nontrivial Collatz cycle."

As I currently understand it, all we know is that:

- a mathematician produced a Lean-verified counterexample to the Collatz conjecture, demonstrating a bug in the kernel

- he claims that LLMs were involved somehow but pointedly refuses to specify how

- he admits that he knew about the bug before publishing the counterexample to his repository.

Perhaps not a joke (although it sure seems to me like they discovered a bug and thought falsely disproving the Collatz conjecture would be a flashy way to announce it), but at best extremely sensationalized by the above description. If you have additional context I would be happy to hear it!

zahlman 2 hours ago||
Indeed. It seems to me much more likely that the AI was directed to look for bugs in Lean, found one, and then it was directed to write a proof specifically targeting the bug.
danielrmay 2 days ago|||
Fascinating, and arguably an illustration of why the bifurcation of responsibility is interesting in the first place.
traes 2 days ago|||
I'm not an expert at it myself, but my understanding is there are numerous ways to "cheat" in a Lean proof (via `sorry` and similar). They're taking responsibility for fully verifying that none of these cheats were used (and that the theorem statements themselves were all correctly formalized.)
rencrisa 54 minutes ago||
Even beyond cheating with sorries or kernel bugs, the lean encoded theorems (or specifications) must be checked by humans to see if they truly mirror the real theorem authentically.
emil-lp 2 days ago||
No, the correctness isn't for the "inside the Lean proofs", but for the translation of "human language math" and its formal Lean variant.
danielrmay 2 days ago||
I see. It still feels like a bit of an oddly solemn way of saying "this is the part we admit responsibility for"
jhanschoo 1 day ago|||
Traditionally, a mathematician would be implicitly responsible for all that (if they were to publish Lean code) and also the intellectual work that led to the artifact of the mathematical paper (and code, if part of the contribution). This statement should rather be read as an acknowledgement of limitation of authorship from the implicit, traditional understanding.
emil-lp 2 days ago||||
Well, to be fair, with Lean proofs, that's the only thing there is (unless I'm missing something).
baq 2 days ago|||
It’s more than you get from free software - you get no proofs, no warranties and any responsibility of its authors are their pure good will. Reminder lean proofs are software!
emil-lp 2 days ago||
I wonder what the total cost of this research was, including the salary for their mathematicians and engineers.
kingstnap 2 days ago||
Why would you factor in salary unless they had to baby it through. You would only count the hours for setting up the harness and prompt and checking the result.

Training the model is going to be amortized over other uses.

emil-lp 2 days ago||
> Why would you factor in salary

Say that it turned out that the total cost of the proof of the Erdős unit-distance conjecture was $50 million.

Then the question really becomes: yes, these models are capable of proving important mathematical results, but at a very high cost. Is it worth it?

If a mathematician applied for a research grant of $50M USD for proving the same thing, they would have been laughed out of the bank.

What's more is that when you have a research grant, you train PhDs and postdocs, you hire new staff, and you disseminate. That is, you get much more value for the money spent.

I'm just curious what the cost is.

ianm218 2 days ago|||
It feels like the real cost might be negative though.. They use frontier math as a way to test improvements in their model. So solving the problems is like a positive externality, but the important thing is they can verify that the model is improving instead of looking at useless benchmarks. Plus it is good for marketing and attracting talent.
kingstnap 2 days ago||||
I didn't argue that knowing the total cost is uninteresting. What I was saying is that realistically the total cost is:

Hours needed for prompt + Hours needed to check result + API costs.

You don't say "well let's add together the total yearly compensation of all the engineers and mathematicians at OpenAI that were involved" and throw that into the total cost. That's simply nonsense accounting.

The actual comparison you are making is some university researcher weighing between getting a grad student (several tens of thousands of dollars) vs typing up a prompt and sending a request to OpenAI for inference (as mentioned in the article, around $2000 in API and maybe a few hours for the prompt and harness).

4fr2 3 hours ago|||
[dead]
z7 2 days ago|||
> The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices.

https://x.com/polynoamial/status/2083470822258467194

AngryData 2 hours ago|||
Yeah sure but they didn't just throw a 5 year old at an LLM with $2,000. If you want good math results from LLMs you need to have math PhDs.
traes 2 days ago|||
That's clearly just for the tokens, this doesn't really answer OP's question.
traes 2 days ago||
Given that OpenAI pays their employees with stock surely a breathtaking number, but not a very meaningful number now that the infrastructure is in place and the models are trained. AI could never get better and it would still be incredibly disruptive.
avaer 2 days ago||
What happens when OpenAI et al stop being open about these things, and just pack it into the training?
traes 2 days ago||
Not much point to pure math being kept secret, in all honesty. There isn't really industrial value, its only purpose (to them) is showing off their model's capabilities. More realistically they'll just stop paying for it.

Edit: Oh, are you suggesting they just use it to privately improve their models? I imagine a few more correct proofs would have a very marginal benefit, if any. Also, they'll probably just get extracted, meaning it still gets out but OpenAI doesn't get to fancily announce it themselves.

asdewqqwer 2 days ago||
At this stage. No doubt calculus had plenty industrial benefit.
simianwords 2 days ago||
What does this even mean lol. These are not solved questions. The solution never existed.
frenzyguy 2 days ago||
This is both awesome and terrifying for mathematicians, however some ideas can be generated and the field as whole expanded with the attention!

However, I was looking at the proofs and reason explanation and openAI should be more explicit in how the work has flown. I find the models have jumped hoops in some places of the proofs, that can be hard to track. In fact, when a paper is published you usually get a review and if no reviewer understands they ask you to further explain the thought process. It will be fun to see if this happens here.

joshlk 2 days ago||
Some of the Lean proofs are 50k lines - is that normal?
qnleigh 1 day ago||
Can anyone comment on the significance of any of these results for their respective fields? Or what impact they might have? Presumably none are quite at the level of the Jacobian conjecture, but some of the results on group theory and sphere packing sound pretty important at first glance.
qnleigh 1 day ago|
Found some discussion here [1] from someone who actually worked on a few of these problems.

[1] https://x.com/henryquantum/status/2083623695436623915?s=20

lifeisstillgood 2 days ago|
On the token limits etc - one assumes that OpenAI et al are able to “hire expert in field, and let them spend the equivalent of a million dollars of tokens” because they are not actually selling their complete compute 24 hrs a day, so the cost internally is a negligible (ish) electricity bill.

Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.

lwansbrough 2 days ago||
For OpenAI, research is marketing. I’m sure they’ve got plenty of budget for that.
paxys 2 days ago|||
No such thing as free, even internally at a company. All such use of resources is accounted for, assigned a dollar value and billed to some department. Someone ran the numbers and figured that whatever they spent on these GPU cycles was worth it.
traes 2 days ago|||
Presumably it's a rounding error compared to their full output, and they're making sure they have enough compute set aside for research by limiting public models. The more datacenters they build the less they have to limit them.
Davidzheng 2 days ago|||
RL training can use all of them - idk what needed means.
simianwords 2 days ago||
I love how people come up with creative ideas to prove the bubble. This one is even more ridiculous - that OpenAI had spare compute to advance mathematics proves that data centres will not be needed. WHAT.

If anything it proves more data centres are needed. That's literally the only reasonable conclusion from this news.

lifeisstillgood 2 days ago||
Sorry I thought that a bubble was widely accepted.

Are you arguing there is not an AI bubble, and that all the DC buildout is fine, going to be profitable etc?

I am not looking for a online slanging match - just looking for a different point of view

simianwords 2 days ago||
It’s not obvious at all. If it were obvious to you, it would’ve been to OpenAI. It’s in their interest to accurately predict demand. The assumption that OpenAI/Sam is both really powerful but simultaneously ignorant to know what others know as obvious is well.. just strange. Especially strange when OpenAI has more information on models, breakthrough and usage patterns and we don’t.

I’m not participating in the slinging match but it’s very very weird that you think it’s some established thing that these companies won’t make profit. A lot of hubris must go in this kind of thought. Like.. do you all think everyone’s playing musical chairs?

AngryData 2 hours ago|||
You say that like the tech world isn't littered in a field of dead and failed companies and billions of dollars burned on failed ventures and ideas. Sure LLMs have proven they have value, but where is the trillion dollars of current investment going to be paid back from? So far it is still entirely speculation that they have such a high value, and they can't just play the long game of "well after a few decades of production and iteration it will add up" because half the hardware cost is going to be obsolete energy burning trash for them in 5 years.
svieira 2 days ago|||
They very often have been in the past. Why do you think this time is different?
simianwords 2 days ago||
“Often” is load bearing. I don’t think markets are more likely than not to be musical chair shaped. To make this conversation more concrete, give me a falsifiable prediction on there existing a bubble. And then I’ll tell you if I believe in it or not.
effseven 2 days ago||
The price of inference is going to go so low that OpenAI and Anthropic will not be able to turn a profit, thus cannot afford the investment into more data centers, thus crash due to investment in the space having been overdone
More comments...