Posted by milkshakes 5 hours ago
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
https://x.com/polynoamial/status/2083478171975082334
As a complete guess, it seems like they tested hundreds to thousands of problems with a relatively low per-problem budget
--
The linked tweet from Noam Brown at OpenAI reads:
> And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).
> But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.
It's not just about requiring to disclose AI use. AI-powered mathematics is a completely valid discipline that doesn't need to be shy, but it should develop its own publication culture.
Pure math is relatively outside my domain, so I find it difficult to grok the exact relevance of the various published discoveries beyond that they are not insignificant, and LLM competence is expanding quite steadily across the field. If this trend continues to the point of LLMs being able to competently expand pure math, it seems somewhat predictable to expect there to be a number of people aiming to find ways to try to keep human mathematicians in the loop.
I've no idea what I think about this one way or the other, beyond that it's certainly a phenomena and one that's going to drive motivated reasoning that may not be entirely sound.
My perspective is more like a FOSS philosophy for math. Even if a closed version has the same immediate effect, it's just better for everyone if everyone can look under the hood and tinker with it.
I would've thought pretty much the exact opposite. "Prompt engineering" was somewhat important in 2023/2024 when the models were much weaker, it doesn't seem at all necessary anymore (unless just "clearly stating your requirements" counts as prompt engineering). Most of the discussion I've seen seems consistent with this?
Methods are only really necessary for results at a meta level, about the design amd evaluation of AI math systems.
shouldnt the paper be the math of the argument? the reproduction is reading the following the proof
Another case I want to highlight is writing GPU kernels as illustrated by the following example: Say I want to generate random number with Normal (0, 1) distribution. Often times the AI written kernel will just generate the number 0. The tests often fail to catch these errors.
Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.
Do you really think that if you paid that to humans, they will deliver the same results?
And frankly these "concerns" ignore reality. In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (and publish it). Openai telling its computer to do that is no different that your phd advisor telling you that.
The post that started this sub-thread asked:
> 1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
I think it's an extremely relevant question to ask, because it helps us better understand the current state of AI being able to handle math, for exactly the reasons I outlined. I was arguing against the idea this is just a reactionary anti-AI kind of question to ask. It's not! You can be very impressed by what AI is capable of in math (I am) and still think those are really interesting things for OpenAI to disclose (I do).
OpenAI specifically called out a $2000 per problem average, which implies something that's probably not true ("if you throw $2k at us we'll solve an open problem for you"). It would be cool to know what the actual number is.
If these 10 problems were solved by humans, it would be pretty impressive, even if it took a large number of researchers! Yet when AI does it, HN commenters suddenly feel the urge to play accountant.
But that's the start of math research, not the end.
The point is to get practice and experience doing research.
Did ChatGPT learn anything from these proofs, that it can build on?
Part of what's annoying people is that ChatGPT is churning though problems that are meant to be motivating. They are problems that aren't worth the effort of human professionals (usually because they are incredibly computation-hevy, so better suited for a computer than a human), so they are good for students to work on.
I know it's more exciting to say "AI disproved a longstanding conjecture" vs to say "it did so AND it took several PhD specialists in the field this many attempts to even produce a prompt that got the model spitting out something useful under some configurations, and many iterations to optimize the configurations, and the prompt itself, and many trials with that configuration to solve the problem. All told we spent more than a typical math academic can hope make in their career."
By not being transparent, they invite skepticism and cynical takes, like maybe it's just that tempered and qualified claims are an existential threat to companies that are fully subsidized by the hype train?
I don't know. Either way, it seems like it would be easy to address these, so why should they not do it?
To be clear, even if that tempered version is close to reality, it doesn't make the models not useful! It just forces a certain calibration of expectations
I say this btw as someone who uses these things extensively, including to disprove an old conjecture my advisor and I were stuck on recently. I know they are powerful and that everything is different now because of them. Let's be sober when discussing them though
That's not normally how people act when they're confident in their product
The cost of running a model is not only $/token, but the salaries of the people managing/orchestrating the models, deciding what theorems to try, etc. Once we factor that in, how much are we really paying per theorem?
The other factor is the subjective component of the value of a theorem. Not all theorems are created equal, and the only way to really measure the value is to ask professional mathematicians for their opinion, or publish the results and look at citations over months/years.
Once we have both of these nailed down, then we can start to do the cost/benefit analysis. To be fair, we should actually compare three groups: human experts, hybrid agent/human expert teams, and fully autonomous agents.
OK I’ll grant that it’s not your obligation to be my search function (despite you making the wild assertion in the first place), so instead can you just point us to the latest grad student solved problem of this level that you know of?
It's a marketing post from a huge company. Only the naive would view it uncritically without assuming it's been written carefully to present the results in the best possible light while skirting the boundaries of outright lying.
Imagine 2 years from now: "yes, GPT solved the Riemann Hypothesis, but cmon, it's just a marketing stunt to hype their stuff, it was probably Terence Tao doing the work but he's so obsessed with hyping AI that he doesn't want to take credit"
We're saying look critically at the claims for how it was done, that it only cost $2000, etc. it would be extremely easy to run 100 sessions that failed, each costing ~$2000, and then just publishing an article about the one that succeeded, for example.
This goes double since it's an internal secret model (Astra) so nobody else can verify the results.
More generally, do you expect that there's some capability threshold where people will no longer study or analyze AI model outputs, and instead just sit there slack jawed saying "so cool!" every time OpenAI announces novel ones? I don't really understand why that would be or why someone would want that. If you're interested in the pure experience of a complex machine outputting satisfying results, I'd recommend getting into sports cars.
Look at their recent claims about their model "escaping" - there was literally a Guardian article calling them out for being hyperbolic! Again, it wasn't that they lied, their marketing department is too savvy for that. They just present it in way that's, well, marketing.
As for the actual result, I'll look for secondary posts by actual mathematicians and draw my conclusions there, not from this marketing blog post about results from a secret model.
Company X does not make money from proving theorems but does make money from selling you a service which supposedly proves theorems. Company X then proves some theorems and explicitly calls out they were very cheap to prove using its service.
And you think you're actually clever for taking these facts at face value? Interesting.
i would classify you as a flat-earther if that happens.
brother like 3 people have pointed out what they're skeptcal of is cost not LLMs - at this point you're willfully misconstruing what people are saying to you just to get a kick out of repeating your same tired strawman.
If OpenAI solved Reimanns hypothesis and the first comment is says something about lack of transparency and marketing, i would say it’s ignorant.
If OpenAI claimed the conjecture to be true but provided no details about the proof then the first comment should absolutely be about lack of transparency.
do you really imagine a scenario where OpenAI would claim to solve it and not give details about the proof? how is this even possible? why would anyone believe them?
That's one of those phrases you can use to dismiss opposing viewpoints without actually engaging with them.
I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.
Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits.
> Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
I think their output has reached a level that precludes this possibility, but I of course don't have any hard proof.
[0]: https://www.reddit.com/r/math/comments/1uxj3cy/after_openais...
I have no affiliation whatsoever with any AI company, nor any formal education outside high school, for what it's worth. Simply being curious and persistent can get you quite far in my anecdotal experience.
https://arxiv.org/html/2605.22763v1
> Our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures
> Our full-featured agent autonomously solved 9 Erdős problems out of 353 attempted, including two questions that had been open for 56 years
Note _had_ been open, not _have_ been open. Can you clarify?
But "had" still doesn't mean what you are implying: once the model solved the problems and the solutions were verified, the problems weren't open any more, so a later description using the past tense is totally consistent.
So I would like to counter your cynicism with a “YMMV” depending on who you work for.
Like cool my lung xray only took minutes to determine if I have a lesion instead of a week or a few days, but I still have cancer.
More importantly, if you can screen for cancer in a way that takes minutes instead of a week, imagine how accessible this technology will become.
But I don't think the argument needs to be "AI is like gambling". The argument only needs to be "humans often behave irrationally and even self-destructively".
Perhaps you can ask Claude to explain it to you.
AI probably did not take your job yet. How many AI queries did you use last month, and how much time has it saved compared to digging through the web?
Imagine humankind meets another race, another race shares it's scientific knowledge and humans accept it without experiencing process of discovery. In that case do we really got this knowledge? If we follow machine discoveries like we follow problems in textbook then we acquire knowledge but we don't discover anything. We follow.
There whole lot of philosophical questions that aren't attacked now. Are complex systems sentient because consciousness is emerging behavior? Then should they have rights? Philosophy is part of humanities and science (is/used to be) part of philosophy. Should we accept non-human knowledge in science? Maybe it's altogether different thing from science, yet very similar.
https://garymarcus.substack.com/p/openais-amazing-but-vastly...
https://garymarcus.substack.com/p/two-critical-updates-re-as...
Not that there isn't something interesting in here, but lets be clear that we don't have enough information to evaluate this properly. And as always with these labs, BS takes a lot more energy to refute than it does to spread.
> Astra, a new model that OpenAI is testing internally, is amazing. No denying that.
To sharpen that, I think he's (obviously) interested in maintaining his own brand as "thought leader" and this necessitates de rigeur defense of particular postures.
Sometimes this is easy because the facts warrant it; other times, a bit of rhetorical license is required to preserve nominal coherence and (at least, for the moment) hold certain lines.
This is one of the latter cases, and it's not subtle.
One of the celebrated properties of many intellectual advances or inventions in whatever domain is precisely that it appears obvious in hindsight. It is quite cynical to leverage consensus distrust of large AI players, warranted but also a popular social construction, to insinuate that these are not "real" advances or "real" hard problems, on the grounds they were in some sense cherry-picked.
Identifying the problems amenable to strategies on the table and intuitions (sic) about where bridges might be, is exactly the discerning work that is the core driver of almost all prior progress, but for celebrated accidents and flashes of insight. Anyone working in any challenging discipline knows that those are celebrated and told around campfires precisely because meaningful durable results arising like that is so uncommon.
These two articles make me think of nothing so much as my own durable reaction to the creeping goalposts of AI critics generally: that they often seem to me not unlike a water color cohort scoffing and jeering at the horse, because it got a D on its tensor calculus exam.
Marcus should be on guard against his own cynicism and take care that his assumptions do not prevent clear sight.
I think that is an oxymoron.
""" 1. By 2029, AI will still be unable to watch a movie and accurately explain the characters, events, conflicts, and motivations.
2. By 2029, AI will still be unable to read a novel and reliably answer questions about its plot, characters, conflicts, and motivations beyond what is stated literally.
3. By 2029, AI will still be unable to work as a competent cook in an unfamiliar kitchen.
4. By 2029, AI will still be unable to reliably create more than 10,000 lines of bug-free code from natural-language instructions or interaction with a nontechnical user, excluding simple assembly of existing libraries.
5. By 2029, AI will still be unable to convert arbitrary mathematical proofs written in natural language into symbolic form suitable for formal verification. """
There's still 3 years to go and he's already wrong on 4 out of 5.
1. still not wrong? Unless it's just feeding the audio or screenplay I don't think you can feed AI a full movie in a single context window yet?
2. Not sure, but can you prove this wrong? Can you feed a full, unseen new book and get that kind of answer?
3. Not wrong.
4. I think he'd probably pull you up on 'bug free' - I don't think that frontier models can reliably write 10k LOC without _any_ bugs typically (not that humans can do this either).
An LLM could theoretically try to earn some money and pay a human to do all 5 tasks but it's clearly not the spirit of the challenge.
Have these been tested or are you just guessing?
Maybe good AI paper writing is further away than I thought...
I'd honestly rather they just automate every job at that point.