Top
Best
New

Posted by milkshakes 4 hours ago

Ten advances in mathematics and theoretical computer science(openai.com)
282 points | 546 comments
sothatsit 42 minutes ago|
People argue whether we are at y-5, y, or y+5, meanwhile we seem to be on a y=2^x exponential that keeps delivering more and more impressive results.

The most interesting question to me is what will be consumed by the exponential like math seems to be undergoing, and what won’t. Writing has been quite stubborn, but I’ve noticed Fable to be quite a big step up there. How about politics? Will we develop new ways to let people express their own values in democracies, or will we just get much better at manipulation? How about experiment driven domains like biology?

tyre 32 minutes ago||
We will get much better at manipulation and better at people “writing” things to justify their own feelings.

What’s new about LLMs is that you can scalably manipulate people individually. It used to be that you could either have scale (speeches, tweets, interviews, website, etc.) or individual engagement (replying to mail/tweets/town hall questions.)

Now you can pull the history and preferences of an individual, then shape a message—in real time—to them, specifically. You can have conversations on social media with a single person and shape your message specifically to them.

Part of this can be good (you talk about what they care about, where 90% of broadcast messaging might not apply) and part of it can be bad (manipulation.)

My guess is that, in the US, the right will cynically adopt manipulation to great effect and the left will take a moral stand against shady practices and lose elections.

lettergram 23 minutes ago||
> My guess is that, in the US, the right will cynically adopt manipulation to great effect and the left will take a moral stand against shady practices and lose elections.

I think that statement may itself highlight how prevalent manipulation is.

I fully anticipate all groups to continue maximal manipulation they can. One thing with LLMs is that it'll be a far less unified view, so a "divide and conquer" strategy is what I anticipate.

jcims 37 minutes ago|||
>Will we develop new ways to let people express their own values in democracies, or will we get much better at manipulation?

Yes.

porridgeraisin 12 minutes ago|||
Today's models depend on inference time compute to get these results. The inference time compute available on any claude subscription is not comparable to the ones used to get some of these results (yes, in this case, it is 2000 USD total as noam confirmed, but some previous results took more).

In general, you can think of the process as generating massive rollouts in generation N, and then compiling in the verifier/human feedback("gradient") signal into generation N+1. The time taken to make the rollout in generation N, and separately the time taken to get the same rollout in generation N+1, each grows constant in some tasks, linear in more, and exponential in some.

In the end, this becomes bottlenecked by time. Today, we can make statements like "I generated all these successful trajectories with 2 weeks of compute, in the next model it will be able to do it in 7 hours of compute", but very soon you'll find yourself making statements like "I generated.... with 8 months of compute, in the next model it can do it in 6 months", which isn't really enticing the same way you can _technically_ brute force passwords but it just needs prohibitive amounts of time and money. That is the "plateau". Note that, this point is quite far away. For example, at any point if we agree it plateaus, today's known hardware techniques such as fixed function accelerators give you a 10x timeline reduction immediately allowing for a few more cycles of improvement. This is not to mention future innovations, but of course none of that is helping with the benchmarks where the time needed is growing superlinearly.

In many math and coding benchmarks, we are still in the constant phase. These are the massive improvements we see every few months. I'm not making any prediction of what will plateau and what will not as it's not possible to make an informed prediction about these things IMO. But the observed fact is that some have already plateaud as in, they don't improve with reasonable inference time (likely superlinear growth).

> will we need mathematicians to translate

Let's take a sudoku analogy. The model is initially just doing the random value algorithm, but lets say you the human are watching it. You make one of the usual reductions and interject "hey you can stop trying 8 here because of ....". Over enough examples, you get to a point where the model is _forced_ to learn the logical pattern. Next generation, it will skip that number. After this, you can peak the distribution using simple 1/0 RL. Doing _pure_ 1/0 RL works decent, but its not frontier as its a very sparse signal.

For that lift, human (or even a better LLM, but if you're trying to improve a frontier LLM, there is by definition no better LLM) feedback becomes necessary. This is _why_ it is crucial that these models interface in natural language and is also why the labs are hiring AI tutors by the hundreds.

> But the long term is completely bewildering if you believe any of these trends can continue at a similar pace for the next few years.

For math and coding, for now we are in the phase where the times are just ... constant, so there's little reason to think it will stop soon. We still need humans to expand the frontier. It just becomes a matter of if its worth the cost of compute for running this generalized The Algorithm or not.

Given how well chess players internalized _many_ (not all) of alphazero's emergent chess knowledge, I am confident we wont have too much trouble figuring out any new math LLMs come up with, which will let us keep expanding the frontier by giving the LLM the next "lift". Only when we reach the stage where the time growth become exponential will this stop, IMO.

viccis 15 minutes ago|||
>Will we develop new ways to let people express their own values in democracies, or will we just get much better at manipulation?

Is there even the tiniest reason to suspect that the people steering this progress will use it for the democratic good of all?

dominotw 16 minutes ago|||
> The most interesting question to me is what will be consumed by the exponential like math seems to be undergoing, and what won’t. Writing has been quite stubborn,

isn't it clearly split between verifiable not verifiable ? what is interesting about that question.

dominotw 18 minutes ago||
> but I’ve noticed Fable to be quite a big step up there

what did you notice ?

DrBazza 2 days ago||
Replace philosophers for mathematicians and Douglas Adams was spot on again.

Whilst current models can't 'intuit' and come up with conjectures, they can certainly disprove some of them very quickly through the kind of grind that humans can't do. I suppose there really are some mathematicians out there today, whose last few years of study, have just been up-ended by this.

--

"Yes we are," insisted Majikthise. "We are quite definitely here as representatives of the Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons, and we want this machine off, and we want it off now!"

"What's the problem?" said Lunkwill.

"I'll tell you what the problem is mate," said Majikthise, "demarcation, that's the problem!"

"We demand," yelled Vroomfondel, "that demarcation may or may not be the problem!"

"You just let the machines get on with the adding up," warned Majikthise, "and we'll take care of the eternal verities thank you very much. You want to check your legal position you do mate. Under law the Quest for Ultimate Truth is quite clearly the inalienable prerogative of your working thinkers. Any bloody machine goes and actually finds it and we're straight out of a job aren't we? I mean what's the use of our sitting up half the night arguing that there may or may not be a God if this machine only goes and gives us his bleeding phone number the next morning?"

WarmWash 2 hours ago||
>Whilst current models can't 'intuit

That's how they are finding these solutions though, unless we are just going to label intuition as something only humans can do. Like a submarine being unable to swim or whatever that example is.

sdenton4 1 hour ago|||
Some of them...

The two places were seeing lots of movement are:

* Updates to lower/upper bounds. In many cases, these kinds of problems are the deep-math equivalent of calculating more digits of pi. Yes, if you throw time at it you'll break the record, but it may not be terribly worthwhile.

* Finding counter examples which disprove conjectures. This is really useful, and helps offset some positivity bias on the human side, often bringing together known tools from distant silos.

If you read the list of ten results, almost all fall into one of these buckets.

pama 1 hour ago|||
It is unfair to dismiss contributions to decades old open problems as equivalent to calculating more digits of pi. It missed the mark by a lot—as does the two bucket simplifaction.
denismenace 43 minutes ago||
How does calculating more digits of pi help us?
tuatoru 30 minutes ago|||
"It's just brute-forcing the search space."
buddhistdude 22 seconds ago||
It can move to any place within the search space but it can't move outside of it and it can't move in between the 'pixels'. Human thought can, as human thought has created the search space.
robotpepi 1 hour ago||||
it could also be that they try every possible approach that has been proposed by humans. it seems that was the case for the non sofic group example. humans are not able to do the same at that scale. it's unfortunate that we don't know what's happening behind the hood with these models, and that's a huge danger also for the rest of us without access to them.
rirze 2 hours ago|||
> "matrices"
fasterik 1 hour ago||
Saying that AI is "matrices" is like saying human cognition is "neurons." Maybe true at some level, but it's a low-level implementation detail. The important part of a language model is the function that maps tokens to contextual embeddings. You could compute this function using analog computing, biological neurons, or any other substrate.
pama 2 hours ago|||
> Whilst current models can't 'intuit' and come up with conjectures

I disagree. I routinely let LLMs speculate or generate hypotheses along the way of helping with technical research. Sometimes they can prove the correctness of a concrete math idea but other times even an unproven conjecture helps with the numerical algorithm implementation and the result is then simply supported by additional data. I guess that any autoresearch-adjacent application has LLMs intuiting and coming up with hypotheses/conjectures—as do the steps/lemmas along a complex proof. In my opinion the modern LLMs are powerful intuitive thinkers that generate lots of conjectures of varying quality or importance.

zahlman 29 minutes ago|||
> they can certainly disprove some of them very quickly through the kind of grind that humans can't do

Of course computers can grind in a way that humans can't. But now we have systems that convert the human-comprehensible ideas into a computer's plan of attack, in a way that greatly expands the frontier of ideas thus treatable.

evenhash 2 hours ago||
> Whilst current models can't 'intuit' and come up with conjectures

People keep saying this. Why?

Surely the AI can complete the prompt “Generate new research questions based on these observations”?

When I read the reasoning traces of coding models they are constantly asking themselves questions and attempting to answer them.

5555watch 1 hour ago|||
I like the illustration that the models are working on a convex hull of known information. Filling gaps with linear combinations of known facts and results.

They can't exit the hull until the "intuition" starts spawning points outside the convex hull.

tuvix 1 hour ago||||
All arguments like this boil down to semantics at a certain point, but yes large language models can “intuit” because they can generalize between examples. The issue then becomes how you pack new examples into context.

Humans can “intuit” based on a much larger, if not unlimited, context. Also I just want to say that human cognition is something so insanely complex and deep that we will not understand it at all in my lifetime. To attribute all, or really any, aspects of human cognition to a machine at this point is silly to me.

michaelmrose 54 minutes ago||
Define insanely complex and deep in a way that isn't illiterate hand waving.

Most humans are dumber than a box of rocks. Here in Seattle we had one of many light rail-related fuckups where they had to replace part of the line with buses. People piled into the front of one when it was full. When people got out they never moved back. As the driver struggled to close the door and people struggled to get in the wad of people never moved back to fill the ample space.

Chatgpt was smarter than the average person a while ago

zahlman 24 minutes ago|||
> People piled into the front of one when it was full. When people got out they never moved back. As the driver struggled to close the door and people struggled to get in the wad of people never moved back to fill the ample space.

This does not demonstrate a lack of intelligence. It demonstrates laziness and a lack of interest in spreading apart. Or just lack of consideration (or even malice) on the part of those at the back of the wad.

> Chatgpt was smarter than the average person a while ago

This is an absurd claim that fundamentally misunderstands what it means to be "smart". Reasoning that would get you to this conclusion would equally well apply to Google's search engine over a decade ago.

tuvix 46 minutes ago|||
I’m not talking about the actions we take or how we might perform at certain tasks, I’m talking about how our brains actually work. My point is that we have no idea how I’m able to imagine an apple and see it in my mind’s eye. It’s basically biological magic to us at this point.

There are processes at work there that we don’t even have the language to describe.

zahlman 22 minutes ago||
Not only that, but we do it with a processor that is basically required to operate in a narrow temperature band below 40C, using a mere 86 billion neurons (although the equivalence with either machine-learning "neurons" or LLM parameters is not at all clear) operating on a few dozen watts; and with this we operate many other systems besides language processing. It's not clear that our reasoning process requires language, either.

(86 billion is the number ChatGPT, ironically enough, has given me a couple of times. I remember hearing for a long time that it was estimated to be somewhere in the ballpark of 100 billion. This is not my field of study.)

claytongulick 2 hours ago|||
> People keep saying this. Why?

For the same reason that you can't draw a 15 of Diamonds from a regular card deck.

treis 2 hours ago||
Of course you can. Tape a 7 and 8 of diamonds together and boom 15 of diamonds
muchmirulys 3 hours ago||
problem number 1 and 9 are surprisingly very intuitive

check here : 1. high dimensional sphere packing https://muchmirul.github.io/conjectures/sphere-packing/

2. multicolor ramsey number https://muchmirul.github.io/conjectures/multicolor-ramsey

dash2 1 hour ago|
The first link is very sloppy and doesn't actually explain why the "certificate" proves anything about the sphere packing. Or if it did, I couldn't understand it.
rothos 1 hour ago|||
Agreed
CGMthrowaway 1 hour ago|||
[dead]
Chance-Device 2 days ago||
Pretty cool. The impact of AI is getting undeniable, there aren’t many positions left to move the goalposts to at this stage, next they’ll have to be outside the stadium entirely.

The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.

fhfncjcc 1 hour ago||
The models are frequently getting worse at items that they aren’t being benchmarked for — and that’s happening more and more over time! Other people in other fields aren’t idiots, they are accurately perceiving the fact that these models are being hyper optimized for our industry, and are becoming less capable in other domains over time. Models of the same scale are massively worse at writing a broad variety of styles of prose than their equivalent from two years ago. (Models of increased scale are a mixed bag.)

Maybe you’re the one who needs breaking out of your cached beliefs.

Chance-Device 5 minutes ago|||
So your answer is: ignore the progress, it’s not really happening, actually it’s getting worse.

That’s not a credible position, but there isn’t anything that I or anyone else can say to someone who simply doesn’t want to believe something.

Legend2440 1 hour ago|||
Proof?

In my experience modern models are better at all tasks than models from two years ago, especially complex multi-step tasks.

alightsoul 1 hour ago||
you are working on coding. they are working on things like "creative writing" remember that gpt 4o was popular among those who had ai as a romantic partnet?
Marha01 4 minutes ago|||
> remember that gpt 4o was popular among those who had ai as a romantic partner

I suspect GPT 5.6 would be even better at it, if given the same sycophantic system prompt and lack of guardrails.

michaelmrose 44 minutes ago||||
sycophancy

It wasn't "better" it was better at kissing your ass which matches what a lot of people want in a partner.

Legend2440 1 hour ago|||
Well that's on purpose lol. OpenAI does not want you falling in love with their chatbot and have been deliberately training it to be less romantic.
vablings 37 minutes ago||
There have been several cases of suicide and self-harm related to 4o, AI psychosis is a real risk and will probably be in the DSM
gste 10 minutes ago|||
People will be broken out of their denial by actual economic growth. That's what this is all meant to be for... I think we might start seeing some surprising numbers.
arenaninja 2 hours ago|||
It's indeed very exciting. I'm looking forward to new advancements/predictions in physics. Preferably as beautifully explained as E = M*c^2
whimsicalism 1 hour ago|||
it is very hard for people to eat crow, as the replies will show
dominotw 14 minutes ago||
Really? has anyone ever claimed that ai will never be able to prove theorems and conjectures ?
slashdave 2 days ago|||
> The sooner people can be broken out of their denial

There is irony here

danparsonson 2 days ago|||
Never understood all this talk about moving goalposts - you understand that's how science works, right? We improve, we learn, we recalibrate our expectations based on what we've learned. If we never "moved the goalposts", we'd be stuck scoring the same goals over and over.
f6v 1 hour ago|||
> Never understood all this talk about moving goalposts - you understand that's how science works, right?

I agree with the parent that we need to acknowledge that we're at a turning point in history. I lived through some of them (internet, ubiquitous personal computing). But it's somewhat difficult to comprehend the impact of this one for many people.

I do biomedical research at one of the top European research institutions. We're very well-funded, but I can clearly see the gap between us (say, top-100) and top-10. I also realize this gap is going to get so much wider unless we invest heavily in AI access (and I'm not so sure I can sell anything more expensive than $20 Claude subscription to the leadership).

I think people having 6-7 figure SOTA AI budgets will move exponentially faster than those who don't. That makes me worried.

So, for me, it's not a question of recalibrating expectations. We're way past that.

NitpickLawyer 2 days ago||||
> We improve, we learn, we recalibrate our expectations based on what we've learned.

That's not what people mean when they say "moving the goalposts". It means that people are adamant that something wasn't important/hard/impressive once the "AI" solves it. And then they come up with another thing that needs to be solved in order to prove it is important/hard/impressive. And once that happens, they do it again. And again. That's what "moving the goalposts" means.

It's also very much not a new phenomenon. It's been happening since the 1980s. As you can see from this quote from GEB by Hofstadter:

> There is a related "Theorem" about progress in AI: once some mental function is programmed, people soon cease to consider it as an essential ingredient of "real thinking". The ineluctable core of intelligence is always in that next thing which hasn't yet been programmed. This "Theorem" was first proposed to me by Larry Tesler, so I call it Tesler's Theorem: "AI is whatever hasn't been done yet."

mag7269 3 hours ago|||
“AGI will only, truly, be achieved when the machine can destroy an industrial-type toilet after downing a Supreme Burrito and a large Baja Blast.”

-Alan Turing (allegedly)

Chance-Device 2 days ago||||
Yes, this is exactly what is meant by “moving the goalposts”. And it’s a fairly well known expression applying wherever people retroactively change their requirements in reaction to those requirements having been met.
danparsonson 1 day ago||
It's almost like I disagree with your use of the phrase in this context, rather than that I don't know the meaning of it.
monktastic1 4 hours ago|||
So you defined your own idiosyncratic version of "moving the goalposts" and used it to rebut his argument, with a condescending "you understand that's how science works, right?"--instead of being honest that it is you who are changing the definition, and not his failure to understand anything.

I don't see how that's any better.

seanhunter 2 hours ago||
He rebutted the argument about moving the goalposts by moving the goalposts. It’s better by dint of sheer bravado.
albedoa 1 hour ago|||
How were any of us meant to know that you were using your own personal and undisclosed definition of a well-established phrase?
danparsonson 1 day ago||||
If it seems like I don't understand the meaning of that very well-known phrase, then clearly I have failed to make my point. I'll try again. And please note that I will use some generalizations to make my point more clearly, rather than because I don't understand nuance; kindly grant me a charitable reading.

In recent years, I have commonly seen the phrase "you're moving the goalposts" deployed by the "it might be sentient" crowd to shoot down the "it's a stochastic parrot" crowd when the latter respond to a new development with "OK but...". In a well-understood field of inquiry, that would be a clear case of goalpost-moving, in the commonly-understood meaning of the phrase where requirements are retroactively changed in response to them having been met. Thank you OP. 'Artificial Intelligence', and indeed intelligence in general, is very much not a well-understood field of inquiry - in fact we don't even have a common agreement about what 'intelligence' is. We are therefore learning as we go (even after all this time!) but making rapid progress in recent years. When rapid progress is made in a poorly-understood field, then how can our definitions and requirements for success not change? This is arguably one of the most pathological development projects ever - what are the requirements? 'It thinks like a human'? What does that mean? And the answer is we don't know what that means, and we're working it out as we go - moving the goalposts. If we didn't move the goalposts, then by definition we already knew exactly where we were headed at the beginning, and we very clearly did not.

Side note that, in case it's not obvious, none of this detracts from how impressive LLMs are. They're a marvel of the modern age, all the problems notwithstanding. However I reserve the right to stay sceptical about their capabilities.

strbean 24 minutes ago|||
> When rapid progress is made in a poorly-understood field, then how can our definitions and requirements for success not change?

It's in how they change, not the fact that they change. The skeptics seem to have secret definitions for intelligence, sentience, consciousness, creativity, etc. that amounts to "a thing only humans have". Often that thing is equivalent to a soul. When yesterday's challenge (LLMs don't have X because they can't do Y!) is met, Y changes but X stays the same. This is not the process by which a field matures, it is a rhetorical technique used by skeptics to avoid honestly stating or confronting their internal definitions. That can be revealed by asking the skeptic the following:

"Forget LLMs. What if we made a completely physically accurate simulation of a human being?"

Many say no, that simulated human being still couldn't have (intelligence, consciousness, sentience, creativity, ...). This reveals that there is a necessary metaphysical component to those attributes, at which point any scientific-minded person will leave the debate.

monktastic1 1 hour ago||||
Thanks for this clarification of your position. The brief answer is:

> If we didn't move the goalposts, then by definition we already knew exactly where we were headed at the beginning, and we very clearly did not.

The criticisms are directed toward people who did clearly act like they knew, not the ones who were honest that they did not know.

Windchaser 3 hours ago|||
> If we didn't move the goalposts, then by definition we already knew exactly where we were headed at the beginning, and we very clearly did not.

To me, the goalposts were already defined by the person you were responding to. "The impact of AI is getting undeniable", so, the goalposts are "the impact of AI". Probably something like "the impact of AI is high, or will be soon".

Note that this does not depend on things like AI sentience or defining "intelligence" more rigorously, it just depends on AI impact.

claytongulick 2 hours ago||||
> And then they come up with another thing that needs to be solved in order to prove it is important/hard/impressive. And once that happens, they do it again. And again. That's what "moving the goalposts" means.

The fundamental argument that I've personally made since the early days of this is that LLMs are not reasoning, in the way that word is commonly understood.

There are lots of reasons why that argument needs to evolve that could certainly appear to be "moving the goalposts", but let's take an example.

A lot of AIs were tripped up by the question "Should I walk or drive 50m to the carwash?" Several folks liked to use that as an example that illustrates that LLMs aren't reasoning, but as the models have been trained on that specific example, it's of course less useful. An AI can mostly nail it now.

So a different example is needed. A new demonstration of how these things fail at basic reasoning a child can do.

Did I move the goalposts? I don't think so. The fundamental argument stays the same. It's not hard to find lots of examples that trip up LLMs, because they are what they are: statistical inference machines. Nothing more and nothing less.

Useful, sure. But also commonly misapplied to areas for which they are inappropriate solutions.

gowld 3 hours ago||||
What you are doing is "motte and bailey".

The motte is "AI useful". The bailey is "Singularity is nigh".

Windchaser 3 hours ago|||
The unified position which many folks deny is "AI is powerful" or, alternatively, "AI will be powerful soon".

(I'm personally still skeptical about this, but I'm being pulled towards accepting it).

"AI is useful" is too low of a bar, and "singularity is nigh" is too high. "AI is on its way to upending society" is about in the middle, and still vastly contentious among laypeople.

enraged_camel 3 hours ago|||
>> The motte is "AI useful". The bailey is "Singularity is nigh".

But there are people like Ed Zitron, frequently posted and cited here, who disagree even with the former.

scotty79 2 hours ago|||
I think Ed Zitron is mentioned just because his name is Zitron. He just repeats ad nauseam opinions concentrated around one simple, very boring pole on a wild and interesting landscape of emerging reality. Anybody could be doing that. A lot of people do that. Yet no other is named Zitron. And that's why I heard name Ed Zitron hundred times. That's how you become a voice of (a part of) the generation. Just have a memorable name and repeat the same opinion over and over that people can flock around comfortably. Content is irrelevant.

Personally I prefer to follow explorers rather than swamp-sitters.

8note 2 hours ago|||
to an extent zitron is saying its not useful, but as a more nuanced opinion, "ai is not cost effective, nor is it improving profits or revenue"
scarmig 2 hours ago||
From a random article I grabbed of his (https://www.wheresyoured.at/subprimeai/):

"it isn't clear whether generative AI actually provides much business value at all"

"cannot seem to find a product that people will pay for, in part because the results are so mediocre"

"Last week, we got our first real, definitive glimpse of what’s around that corner that future. And boy, was it underwhelming."

"OpenAI claims that o1 “performs similarly to PhD students on challenging benchmark tasks in physics, chemistry, and biology.” Just not in geography, it seems. Or basic elementary-level English language tests. Or math. Or programming. "

"Worse still, it's kind of hard to explain why anybody should give a shit about o1."

"o1 shows that OpenAI is both desperate and out of ideas."

"the software is not becoming more useful"

Honestly, every other line is quotable in this context.

lackoftactics 1 hour ago|||
Yep, he is a PR stunt guy, and the number of videos that come up when you type Ed Zitron into YouTube should tell you how many people are eager to feed their cognitive biases.
dwaltrip 1 hour ago|||
It’s comical and honestly incredibly embarrassing…

But it seems we have somehow optimized away shame. It wasn’t good for profits, I guess.

emceestork 1 day ago|||
They aren't claiming that science doesn't progress by moving goal posts. They're talking about how critics of AI have claimed it isn't revolutionary/useful, then progressively changed what would it mean for AI to be actually revolutionary/useful.

Not long ago many folks were saying AI was the same as the crypto bubble. No real useful technology and only hype.

gowld 3 hours ago||
Did you know that "revolutationary" is not equivalent to "useful", and "revolutionary" is quite ambiguous?
emceestork 1 hour ago||
I don't know if you're trying to dunk on me. I didn't intend to imply they are synonyms.

I think AI is clearly both revolutionary and useful. Revolutionary insofar as the job I do has changed almost completely in a year or so span.

matsemann 1 hour ago|||
Which straw man are you arguing against?
Dig1t 2 minutes ago||
I don't think this is a straw man, a huge number of people in my life (non CS people) think AI is a dead-end, that it's just a stochastic parrot, that it'll never be able to do many things that humans can do. I have had many arguments with people who told me that "AI will never be able to do X", and then 6 months later AI is able to do X. Then they will move the goal posts and say "well AI will definitely never be able to do Y".
c7b 3 hours ago|||
And what does taking it seriously entail?
Chance-Device 2 hours ago|||
In the near term handling the transition. Jobs will be lost, careers ended, people won’t be able to reskill quickly enough. At the same time AI is an enormous opportunity to uplift living standards, but nobody has the logistics of this figured out.

We need to figure out how to restructure the global economy. How does UBI work internationally, if the AI companies are taking revenue in the US? What’s the tax base for it? What does that say about international trade and protectionism? Do countries end up splitting into different trading blocks based on their level of access and legality of AI (I assume some will ban it outright)?.

How does intellectual property work in an AI generated future? What about healthcare advances, who gets to own those?

What about meaning, what about purpose? How do we replace the work ethic that tells us we are our jobs and idleness is immoral? How do you replace “What do you do?” As one of the first questions you ask a new person?

That sort of thing.

bubblemoth 1 hour ago|||
What pressure is there to push for any of these changes? I see AI advocates discussing the concept of UBI, but I can't imagine a world where the United States would ever pass this sort of legislation. I mean, congress can barely pass a budget each year.

If you are correct, I expect corporations to reap massive profits while most Americans try to find a way to survive in a world where they are obsolete.

Chance-Device 1 hour ago|||
There is whatever pressure you bring to bear on it. As long as you are still living in a democracy your vote is the pressure you can apply. If the parties that exist won’t represent you, make new ones that will. Do something rather than deciding it’s both impossible and up to other people anyway.
azinman2 1 hour ago|||
I also don’t understand why UBI is desirable. Putting everyone on welfare means everyone is poor. This won’t end well, and it’s certainly not the case that 99% of the population will let a tiny number of people remove their income in favor of pennies for all. Political violence will come first, easily.
striking 2 hours ago||||
I think this too is a kind of denial. As in, while some are in denial about the usefulness of AI, others are in denial about whose living standards are actually going to be uplifted.

And it's sad, really, because I think these two groups would make a great pairing if they could stop arguing against one another for a moment. They'll both be impacted about as much and probably have the same ultimate goals (to lead dignified lives).

But it seems these days everyone is more interested in Kayfabe and feeling like they're in the right than working together, so maybe I should just keep quiet rather than attract the ire of both groups...

throwaway0123_5 2 hours ago|||
> others are in denial about whose living standards are actually going to be uplifted.

I don't know if it is fair to say they're in denial. For my part, I don't expect life to get much better for regular people (especially short term), but that doesn't mean we shouldn't work to try to make it happen.

Chance-Device 2 hours ago|||
Asking questions about policy and values and pushing to have those resolved in positive ways is about as far away from denial as you can get. It’s possibly the only useful thing an ordinary person can do.

What a lot of people want to do, and I’m not saying that you’re one of them, is to assume that a positive outcome is impossible and either do nothing or loudly yell that the world is ending. Neither is particularly useful.

Or, as I said above, others just deny that there’s anything to see here and try to get people to move along.

striking 19 minutes ago|||
Asking politely is not how we got a 40-hour work week or workers' comp or most other labor standards we take for granted, just historically speaking. I think a positive outcome is very likely, and I think it will be a lot more work than loudly yelling, but I don't think anything will happen if we try to build everything up from first principles instead of taking a moment to consult history as many are wont to do in this AI era.
HarHarVeryFunny 1 hour ago|||
The AI companies themselves, who are highly motivated to sell AI as overall positive for society, notwithstanding some security/etc risks, and who have economists on the payroll to think about things like this, do not seem to have found any possible positive outcome to present.

Shane Legg (DeepMind co-founder), one of the more intelligent and thoughtful people you'll find in the industry, could only offer "it's a tough problem - we need to think about it" when recently interviewed by Hannah Fry.

On the surface the most likely outcome for AI allowed to replace jobs is extraordinarily negative, especially since it is a general capability technology, not a specific one where displaced workers can just move to another field. Once AI becomes more capable it will be able to do the vast majority of white collar jobs, including any new ones that may appear as a result of AI. As Shane Legg put it, "if your job can be done remotely, sitting in front of a computer, then it can probably be replaced by AI".

Not only does AI threaten to replace ALL the white collar jobs, but it is rapidly going after blue collar (factory jobs, driving jobs) and pink collar ones (Japanese robotics for elder-care) as well.

If a positive outcome (which doesn't include putting displaced workers on welfare - UBI) is possible, then it sure would be nice to hear it, and the silence from the AI companies, and government for that matter, is deafening.

Chance-Device 33 minutes ago|||
UBI probably is the positive outcome, though it may not seem like it to begin with. Initially it will likely be stigmatised and under-resourced, but as a larger proportion of people move out of work and onto UBI that stigma will drop and the resources should grow.

Eventually UBI will be the norm, and if the living standards of a person on UBI is as good as yours or mine today, that will be an enormous win for everyone. It’s like pensions, once these were only for the elderly poor, now they’re a right for everyone in most developed countries.

It’s also interesting that for most of human history leisure time was the point of life, and only in recent modernity has work come to be the meaning of someone’s existence.

UBI has to be commensurate with production being automated. That’s a big logistical problem, if you think building datacenters is a challenge try bringing about radical abundance, but even so it’s not insurmountable. It just needs to be taken on as project and not seen as an impossibility.

So much of this is not about what is possible so much as what people believe is possible. We can do anything if we try.

orangecat 38 minutes ago|||
do not seem to have found any possible positive outcome to present

See "Machines of Loving Grace" by Dario Amodei: https://darioamodei.com/essay/machines-of-loving-grace.

HarHarVeryFunny 23 minutes ago||
This seems more a fantasy than a considered likely outcome. He basically admits that humans will eventually mostly all be out of a job, replaced by AI, but then says (Gemini's summary) that there will be:

"Massive Economic Abundance: Because AI will exponentially grow the total economic pie, overall resource scarcity will diminish. The fundamental challenge shifts from producing wealth to distributing wealth."

So how do we go from everyone out of work, no income to spend on food, or the goods and services that the AI is producing, to "massive economic abundance"?!

It's like the meme:

Step 1: Create AI

Step 2: AI takes all the jobs

Step 3: ???

Step 4: Profit! (massive economic abundance)

What is step 3?

cautiouscat 13 minutes ago||||
I’ll be the first to call myself cynical.

> In the near term handling the transition. Jobs will be lost, careers ended, people won’t be able to reskill quickly enough. At the same time AI is an enormous opportunity to uplift living standards, but nobody has the logistics of this figured out.

> We need to figure out how to restructure the global economy. How does UBI work internationally, if the AI companies are taking revenue in the US? What’s the tax base for it? What does that say about international trade and protectionism? Do countries end up splitting into different trading blocks based on their level of access and legality of AI (I assume some will ban it outright)?.

UBI in the United States is never going to happen in time. If it happens at all. We don’t even get universal healthcare. I think people who think AI will be a net positive for humanity are also in some sort of denial.

In a different US political climate I would entertain it. If these frontier labs weren’t so clearly going after the money, I would entertain it.

LLMs are clearly a step up for capitalists so I just can’t see any inclusion of LLMs move towards more progressive ideologies.

thuuuomas 16 minutes ago||||
Why should IP persist in a world of “intelligence too cheap to meter”?
joshmarlow 2 hours ago||||
I don't understand why this got downvotes - simply extrapolating current trends leads to the need to answer all of these questions.

My own $0.02 on the economics piece - every country should have a sovereign wealth fund. Governments should block market access from automated[0] companies until those companies provide equity contributions to the wealth fund for that country. This aligns regulator and corporate interests. Dividends flow into the sovereign wealth funds and then can be allocated locally from there - UBI, job programs, etc. Let different jurisdictions explore different ways to structure a post-labor society.

On the broader social front - I think a lot of lack of meaning discussion boils down to the overemphasis we have on your job as your self-worth. We need to realign our societal expectations - and people need to spend more time with their families.

[0] for this to work, I think we would need well accepted metrics for 'how automated' a company is - and that probably needs a 3rd party auditing industry.

GPerson 2 hours ago||||
Why do you think the AI companies are going to let it be a “we” kind of decision, and since it’s obviously not going to be a “we” kind of decision, as getting to this point certainly has not been thanks to people like you, what makes you think living standards are going to be broadly uplifted?
esafak 2 hours ago||||
Isn't it obvious if AI automates your job away, the AI company is going to reap the economic value, which is going to be less than what you cost, while giving you a pittance as UBI? If you received its full value there wouldn't be any point in anybody replacing you with AI.

The only way to win is to wield the AI.

unfitted2545 1 hour ago|||
Nationalised LLM? As long as the state doesn't decide what information the LLM shares (from an output and privacy perspective).
throwaway0123_5 2 hours ago|||
Agreed, a LOT of UBI advocates gloss over the "B" in UBI. If AI increases human productivity overall, the only morally acceptable outcomes (imo) are that everyone's standard of living increases (or at least is the same without having to work) and wealth inequality decreases (if AI is doing ~all the work, there really isn't any sensible justification for some people having significantly more wealth than others). Frankly anything else seems like a recipe for massive social instability.
hansmayer 1 hour ago||||
[dead]
GolfPopper 2 hours ago|||
Sorry, all the billionaires and their pyramid architects are too busy with the Install Planetary Overlord planning and logistics to think about any of that. But don't worry! Once they're done, I'm sure they'll put their slave super-intelligences to work figuring out how to humanely come up with a final solution to the challenges posed by the remaining uncontrolled humans.
WarmWash 2 hours ago||||
The human zoo where the top ~250,000k humans live in a "human utopia" and the AI provides while mostly focusing on whatever it decides it's own goals are.

Humanity survives (but we reading this probably don't), the AI treats the living humans like the Emperor's favorite pets (probably a pretty good life), and then the AI does whatever else it deems important.

mofeien 33 minutes ago|||
To what kind of goal that an ASI might decide to pursue would "a quarter billion happy, healthy, free people" be the most efficient solution to?
WarmWash 2 minutes ago||
My mistake, I meant 250k*
waffletower 1 hour ago|||
[dead]
logicchains 2 hours ago||||
Realistically it means trying to start a small business of some sort, because AI is hugely advantageous to business owners and disadvantageous to workers. And it's something AIs can't do unless they get legal personhood, which may well not happen any time soon.
GPerson 2 hours ago||
Hopefully doesn’t ever happen, since that’s one of the more plausible omnicide scenarios. I human like AI is almost certainly achievable without much research effort at this point, but we shouldn’t do it.
twister2920 2 hours ago||||
[dead]
kypro 3 hours ago|||
[flagged]
addaon 2 hours ago|||
> "probably only 20% chance we all die"

Bad news for you -- there's a 100% chance we all die. Sorry to be the one to tell you.

cubefox 2 hours ago||
Logical mistake. There is a difference between every human dying eventually and humanity going extinct. The former doesn't imply the latter. The previous commenter clearly meant the latter.
kaonwarb 2 hours ago||||
What makes you think separate nations would effectively cooperate in this way?
armchairhacker 2 hours ago||||
You’re not necessarily wrong, but math proofs aren’t going to unify today’s brainrotted population.
lkey 2 hours ago||||
Your 'serious' proposal is unlimited global military bombing campaign on civilian infrastructure by the United States (which is currently losing a war using the same strategy) to 'solve safety' preemptively against a 20% number you just made up?

And you accuse the 'other side' of 'suicidal apathy'??

You should put down the AI and do some self-reflection on how you came to hold these views.

armchairhacker 2 hours ago|||
GP said cooperation, never implied the US would do it alone.
deaton 2 hours ago|||
If the thesis that "If anybody builds it, everyone dies" is true, or has any chance of becoming true, then it is the logical thing to do. As Geoffrey Hinton said, "If you want to know what it's like not to be the apex intelligence, ask a chicken."
lkey 1 hour ago|||
It is not 'logical' to imagine a hypothetical doomsday scenario that justifies a preemptive nuclear war. (which is what the grandparent commenter's bio contemplates, 'nuke the datacenters' is their central credo).

Does it bother you that the people who are publicly cocksure that P(doom) is moments away are the same people that have profited most handsomely from that pronouncement?

That the 'humanists' that want to do 'altruism' for 'potential future humans' and are the same people that commit fraud and theft at a civilizational scale, then sell this 'intelligence' to any child-incinerating militaries with spare cash?

It's not wrong to want to do good, but if a system that is branded 'do (the most) good' commits great evils, you are morally and intellectually obligated to step back and reconsider how you are spending your time.

Also, I asked a chicken and a feral rock dove what it's like to be not be 'apex' and they burbled at me and kept eating millet and sunflower seeds.

Would you like me to follow up with them? I'm not sure what point you expected them to make.

cubefox 1 hour ago||
> It is not 'logical' to imagine a hypothetical doomsday scenario that justifies a preemptive nuclear war

You hallucinated the "preemptive nuclear war". He didn't say anything about nukes. That's your own invention.

> Does it bother you that the people who are publicly cocksure that P(doom) is moments away

20% is not "cocksure". The "moments" is again an exaggeration.

lkey 20 minutes ago||
I did not, I read kypro's (the OP I was replying to) bio, to whit:

Every problem is a search problem. Nuke the data centers.

P(doom) = 98.9% (Aug-2026) P(doom) = 98.2% (July-2026) P(doom) = 98.2% (Jun-2026) P(doom) = 98.5% (May-2026) P(doom) = 98.8% (mid-April-2026) P(doom) = 98.7% (April-2026) P(doom) = 98.7% (March-2026) P(doom) = 98.5% (mid-Feb-2026) P(doom) = 97% (Feb-2026) P(doom) = 94% (Jan-2026) P(doom) = 93% (Dec-2025) P(doom) = 95% (July-2025)

samatman 22 minutes ago|||
If the Judeo-Bolshevist theses were true, then Operation Barbarossa was the logical thing to do, and everything which came with it.

Powerful word, `if`. "You're not only wrong you're a fulminating psychopath" is a perfectly valid response to getting it wrong like a fulminating psychopath.

DarmokJalad1701 2 hours ago||||
No.
bigyabai 2 hours ago||||
> We are now at the point where RSI is feasible

  What can be asserted without evidence can also be dismissed without evidence.
- Hitchen's Razor
apetresc 2 hours ago||
The evidence is abundant and nearly impossible to ignore without increasingly focused effort.
ghjkghjkghj 3 hours ago|||
There are some genuinely insane takes in this thread on both sides but this takes the cake.
iwontberude 2 hours ago|||
That's because its satire, the tell was "not excluding targeted military strikes"
ghjkghjkghj 2 hours ago||
Gonna say the edit they made removes the satire possibility.
cubefox 2 hours ago|||
A few years ago the majority of Hacker News dismissed LLMs as not much more than stochastic parrots who were decades away from doing any serious work, autonomously escaping containment and hacking Hugging Face, or outperforming most professional mathematicians at proving theorems. These things were labeled "insane" and "science fiction" despite the rapid progress we had already seen.

Now you are again postulating that there would be no more extreme progress in the near future. That's actually more "insane".

GPerson 2 hours ago||
Being right about the thing that’s genocides all culture and likely worse isn’t such a great thing to brag about, but go ahead.
ck2 3 hours ago|||
"AI" has limits in that it cannot invent knowledge, it can only distill and search for patterns in existing knowledge

not sure how many will get this reference but "AI" for science and math is like super-shoes for runners

at first we are blown away by the impossible improvements including sub-2-hour realworld marathon and every other PR/CR/WR is dialed down

but then the improvements slow and reach a stall point because of the limit of technology and the source of the achievement

ie. sub-2-hour marathon yes, sub-1-hour never happening (rollerblade inline-skate record is 1-hour marathon)

fixedpointsnake 2 hours ago|||
I agree. The most likely scenario is that this is just a "new normal" lift that is percolating through human endeavors and will saturate at some point. For example, the whole cyber-security bruhaha should ultimately resolve into higher standards for code published -- we can now cheaply find and fix a whole slew of minor bugs that weren't worth our time before.

The fact we see a lift is not the same as evidence that the lift is unbounded.

The lift being finite is supported by the fact improvements have come at the edges: improvements from human feedback, improvements in harnesses, improvements on model compatibility with harnesses, improvements in inference efficiency with new architectures, etc. If we were just training better models from scratch that would be one thing, but we are just making better use of a tool we've developed.

pama 3 hours ago||||
Not sure what your first sentence means, or why you are quoting AI. Many of these problems individually were math at a level approaching the highest possible for expert human mathematicians. These are not simple combinations of existing ideas, or following of human intuitions, or implementing something following specific human instructions. Then again maybe you mean that math is not knowledge and that all math is simply extending the basic axioms using known patterns, to which I would not agree.
lackoftactics 1 hour ago||||
as much as I love analogy with humans using super-shoes, not everybody can be a world-class expert in their industry. There should be a place at the table for average people to take part; otherwise, it won't be sustainable.

As a programmer, I am mostly interested in whether my role is sustainable long-term and whether the models will get better. I don't feel in jeopardy yet, but two more years like this and the calculus of hiring software engineers could shift even further. QAs are already overwhelmed with work

IncreasePosts 2 hours ago|||
How do humans "invent knowledge"? Is your argument that the answer to these questions already existed in the training set? Why didn't any human recognize that before?
hansmayer 1 hour ago|||
[dead]
datakan 2 days ago|||
[flagged]
Chance-Device 2 days ago|||
I can deal with apathy, that’s the norm. What bothers me are all the people who think they can suppress AI by talking it down. That’s what’s counterproductive, just pretend the problem doesn’t exist. Tell other people it doesn’t exist either. I get it, it’s threatening socially, economically, maybe existentially. It’s also not going away.
ryan_n 3 hours ago||
So you think it’s a potentially existential threat but are bothered by people who maybe want to suppress it… Hopefully you acknowledge there is a bit of lack of self awareness here eh?
FranzFerdiNaN 2 days ago|||
It’s not apathy. It’s the fact that almost nobody can really understand what these results mean.

I’m not a mathematician so I have zero clue what “ New upper bounds on sphere-packing density down to the Cohn–Elkies thresholds” means.

overgard 1 hour ago|||
https://garymarcus.substack.com/p/openais-amazing-but-vastly...

https://garymarcus.substack.com/p/two-critical-updates-re-as...

As always, PR hype. Goalposts have not moved.

Guys, please use critical thinking. The haters don't hate by default, we hate because we're gaslit about this stuff every day and it's annoying. Extraordinary claims require proof, and they're not giving us information that would be essential to knowing if this is actually significant or not.

w4yai 1 hour ago|||
AI already have an impact, and yes this is PR hype because this is a product. Yet both can be true at the same time. We're not blindly eating what's OpenAI is serving us as gold truth, we're just admitting it's doing remarkable progress.

Remember October 2024 Pelicans [1] ? It's been only less than 2 years.

We don't know what will come in the next 2 years. But the progress doesn't seem to stop for now.

[1] https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/

skydhash 53 minutes ago||
The pelicans are still not ok to this day.

People are skeptical of the announcement because the room include several PHDs in math and physics. The prompts are not published so we can see how generic the starting prompt is.

afro88 1 hour ago||||
Gary doesn't argue it's hype though. He argues 2 things: other people are getting carried away with the result, and we don't know enough about how it was reached to know where it falls on the impressive scale.

He literally says it's an impressive feat in the second article.

efavdb 1 hour ago||||
best point in that first link: openai doesn't tell us if they only tried to solve these 10 and each was solved (amazing) or if they asked it to solve a million problems and it got these 10. Either is great, one is more so.
hgoel 1 hour ago||||
Your comment seems entirely disconnected from the posts you linked. It's impossible to deny the results, it is not PR hype that in the past couple of weeks LLMs have resolved problems that have been open in mathematics for many years. Some of those problems had remained unresolved despite keen interest from many humans.

The only way that is PR hype is if you're invoking the insane conspiracy that frontier AI labs are just buying off results that would otherwise be career defining for a mathematician, just for marketing.

The posts you linked are urging caution regarding the exaggerated e/acc-esque lies peddled by people like Musk, not that the models haven't proven themselves as having genuine ability to contribute to research in some areas.

dwaltrip 1 hour ago||||
I use these strange machines all the time. They have gotten notably smarter. That’s my personal experience.

They still do things that I find incredibly annoying and “dumb”. And I still have to clean up messes they make quite often.

But on the whole they are clearly smarter than before. No extraordinary claims needed. I just try to learn how the tool works and how to use it effectively.

bluerooibos 1 hour ago||||
Gary Marcus has been moving the goalposts since day 1. The guy is a psychologist. Why would anyone care what a psychologist has to say about AI? He's likely made good money from constantly moving the goalposts and being a denier, due to the publicity he gets.
bonoboTP 1 hour ago|||
He's simply a good phone number to have for journalists under time pressure who need to add the contrarian voice to their upcoming story. He delivers it reliably, then never reflects on how he was wrong in the past, just blasts forward as if nothing happened and just makes the next bonkers claims to the journalists who are very thankful for the prompt delivery of how AI is a nothingburger, and fake and won't ever do XYZ that it then proceeds to do in N months.

I remember the time when he insisted that diffusion-based image generators trained on Internet scale data will never be able to make an image of a horse riding an astronaut. Today you can generate 4K video of that.

mef51 1 hour ago|||
Because he's not talking about AI, he's talking about people's psychological reactions to AI
Trasmatta 1 hour ago|||
Society at large is getting worse at critical thinking, because we are increasingly offloading that thinking to AI
HarHarVeryFunny 2 hours ago|||
Sorry to hear you've been impacted by this AI math.

I heard that Gary Kasparov was impacted by AI chess, but at least he still seems to have a job, so don't give up.

zahlman 39 seconds ago||
Making insulting assumptions about the hidden motivations of others is not the level of discourse I come to HN for.
applicative 1 hour ago||
It is a fact of experience, and indeed effectively a theorem, that the better they get at coding and math, the dumber they are. These are the wages of RLVR etc
whimsicalism 1 hour ago||
completely false
kcexn 2 days ago||
Not being an expert in any of the fields OpenAI has "advanced" I don't want to prematurely downplay the significance of this contribution. However, I am worried that the language they are using in this blog post is exaggerating for the sake of marketing.

It is true there hasn't been a reliable computational approach to solving these problems before. But do these proofs contribute new ideas to the mathematical corpus, or are they simply an effective method to exhaustively search the literature for the right combination of existing tools to apply to the problem?

Essentially, did these problems seem like they had an intuitive answer and were feasible to prove before, just not high enough value targets for an expert to invest time into? Or were they fundamentally difficult prior to this point and it appears that AI has done something more than just throw the problem into a big solver.

x0mej 3 minutes ago||
You’ve received the expert answer several times. You just don’t seem to like the answer.
Ar-Curunir 2 days ago|||
The problems from CS (CVP and circuit complexity) are very important problems that have been worked on by top researchers for 30-40 years. Some of these researchers include Turing Award winners. A solution to them would be a best-paper award at many top CS conferences.
kcexn 1 day ago||
I assume you're talking about No. 5, the arithmetic circuit complexity bound? The existence of a lower bound than state-of-the-art is certainly a significant result and worth publishing.

But the wording of the result makes it sound like we don't know what the lowest possible complexity bound might be. So, prior to this result did we think there couldn't be a lower possible bound? Or did the arithmetic circuit community think there were lower possible bounds but didn't see it as a high value target for experts to tackle (maybe a problem that was instead regularly given to students to study).

Ar-Curunir 2 hours ago||
Circuit complexity lower bounds (and lower bounds in general) are notoriously difficult to come across.

For example, despite our best efforts, the state of the art lower bounds on time complexity of algorithms for solving 3SAT is O(n). In contrast, our best algorithms for the task run in time roughly O(2^n). That’s an exponential gap. This is despite decades of trying to find lower bounds.

QuesnayJr 2 days ago|||
The ones I'm familiar with are big breakthroughs, but they are both counterexamples. Examples have an advantage in that once you have the example in hand and a sketch of the proof (which they have provided), then an expert can probably work out the details themselves.

The sofic groups question was the outstanding question about sofic groups. Almost everyone thought that non-sofic groups existed, and there were plausible candidates, but proving a group was non-sofic was out of reach. Now that we know how to do it once, we can probably do it a lot more.

The Connes rigidity conjecture I think people thought was false, but it was a provocative claim to make. The significance of conjectures is frequently not that the answer to the question is "yes", but that we don't know how to answer the question. And now, apparently, we do.

robotpepi 1 hour ago|||
> but proving a group was non-sofic was out of reach

a colleague was telling me that the base idea for proving that something is not sofic already appeared in the literature around 2019 or so (this is the "expanders graphs" that are mentioned in OpenAI s paper. no one had managed to find a concrete example though. this doesn't make the result less impressive in any case.

kcexn 1 day ago|||
Interesting. Do you have any more specific insights into where you feel AI was a big value-add to these problems? I don't want to be overly dismissive of AI, but I also feel that the AI hype engine frequently positions claims as being 'ground-breaking' when they are really just interesting incremental results.

The general consensus of developers is that AI can only do the work of a strong 'junior'. Yet as soon as we are presented with pure mathematical results, people seem incredibly ready to accept that AI can do more than what a strong student could achieve.

QuesnayJr 1 day ago||
They are more than a strong student could achieve. I'm not equally familiar with the problems, but the ones I'm familiar with, if a student solved them people would be thinking "that's someone on track to win the Fields Medal one day".

If it works better here than for programming, then I would guess it's because you can give it a very precise prompt, so you either solve the problem or you don't. If you read the prompts people have shared for problems like this, then the instructions are basically "Solve this problem. Don't give up early. Don't solve a similar problem."

simianwords 2 days ago|||
> However, I am worried that the language they are using in this blog post is exaggerating for the sake of marketing

Your worry.... is because they used the word advanced? For marketing? The word is used very appropriately here. There were PhD's who spent a big part of their career tackling these problems.

kcexn 1 day ago||
I have no idea how many PhD's have spent how much time of their careers tackling these very specific problems, and I doubt you do either.

I'm trying to understand if these specific problems were the kinds of problems that would have justified an expert investing weeks or months to solve. Or if they were the kinds of problems that would normally have been given to students to investigate.

hollowcelery 2 hours ago||
They are significant problems which experts have spent months or years studying. I heard a mathematician say that resolving non-sofic groups and Connes's rigidity would be career-defining for a mathematician.
patcon 2 days ago||
"Breakthrough research" can be defined (in the citation record) as research that both (1) becomes highly cited, and (2) brings together citation chains that were previously not showing up together.

Mundane incremental research is cobbled from existing citations that already appear nearby in the record.

Basically, innovative research is a measure of bridging thought and domains that were previously not bridged. It's quite concrete as a measure in the citation record.

So we can know pretty conclusively.

Puja Ohlhaver gave a talk on this[1], and ran some experiments (that I had the pleasure to support on)

[1]: https://www.youtube.com/watch?v=guLDNMAOn24

kcexn 1 day ago|||
I'm not arguing that this isn't innovative or worthy of publication. Basically any result that moves the needle meets those criteria. I'm interested in how the results that OpenAI has published here differs from finding optimality solutions for incredibly niche optimization problems by throwing the problem in an enormous solver.
casey2 1 day ago|||
Breakthrough math research is very rarely highly cited. Maybe some combination of pretraining scale, inference speed and orchestration will help, but it's telling that OpenAI is solving random math research problems rather than bedrock algorithms and their implementation. Even as cool as the tech is, there still is very much a clock that they have to outrace before they collapse.
simonw 2 days ago||
The GitHub repo with the Lean formalizations just came out a couple of hours ago: https://github.com/openai/ten-proofs

It also links to a paper written by an LLM where the model "reconstructs how the proof came together" based on the unpublished reasoning traces: https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf

I wish they'd publish the prompts though!

derbOac 2 hours ago||
I like the Lean formalizations — I hadn't thought seriously of asking for that before but might try it with some stuff I've been working on.
fooker 1 day ago|||
Exact prompts haven't mattered for about a year now.
Alifatisk 1 day ago||
Care to elaborate? Curious about this. Is this because LLMs have been geared towards understanding user user intent behind a prompt rather than following the instructions exactly?
fooker 1 day ago||
There's a full fledged 'reasoning' step that basically expands your prompt.

As long as you are not missing important information, how you word the prompt does not have any effect.

s4i 1 hour ago|||
Isn’t that a huge simplification? Of course the way you phrase the prompt can carry semantic meaning, maybe subtly, but still. And sometimes that matters a little and sometimes a lot. I’ve stopped numerous agent sessions over the last few weeks to reword my initial prompt to get the agent off an unintended track.
Alifatisk 5 hours ago|||
Oh yeah, I suspected it was something like this. Thanks!
hacklewoodple 5 minutes ago||
[dead]
raver1975 41 minutes ago||
I wish I could qualify for some free AI as a mathematics researcher. I guess I'm just an amateur. https://alethean.org
ultimatefan1 2 days ago||
one of the early premises of how ai takeoff would go was that a system that could solve open problems in advanced mathematics would also discover novel advances in math and computer science that directly unlock drastically better software performance. we are seeing frontier level math breakthroughs (ie performance that would put it in the top 100 or 1000 mathematicians in the world if it were a human, meaning top .00001% or 800/8B). we are also seeing incredible advances in software performance. open ai announced like 15% improvement by fixing gpu kernel issues. these are clearly linked in the sense of scaling laws and generalization of intelligence: a huge model gets capabilities in both math and software engineering that isn't possible at smaller scales.

but it seems less likely to me than before that the types of math/science discoveries will explicitly unlock better software performance. in some sense this fits our intuitions. when top tech companies use math PhD type employees, they have them stop doing pure math research and instead focus on software engineering. these people are often very good at software engineering but not due to recent discoveries in academic mathematics, it's due to their general intelligence. to me, this is evidence that the models are getting better but does not make me think we are on the cusp of a foom style fast takeoff enabled by revolutions in frontier math (i also posted this on twitter @mlipman13)

GPerson 2 hours ago||
I don’t really like AI but let’s stop kidding ourselves, no human mathematician could make progress on a dozen major open problems in a week or two. If you’re measuring it against humans then it is by far the best mathematician to ever live.
paulmist 1 hour ago|||
Correct me if I'm wrong, but all of the aforementioned advances were made in the last year? Until very recently few people had access to these tools. Most people still don't know how to use ChatGPT, and very few use tools like CC regularily. If in a few years these frontier tools become commonplace and people upskill we would should see a network effect?
woeirua 2 days ago|||
This makes no sense. To believe this you have to think that the models are somehow being overfit explicitly on academic mathematics and it doesn’t carry over at all to more practical software engineering. I wouldn’t make that bet.
jvanderbot 2 hours ago|||
Or, that the mathematical formalisms that model the limits of software performance are firm enough that barring P==NP, nothing much will change despite proofs of beautiful math.
threatofrain 2 days ago||||
This also makes the assumption that frontier math has all the long hanging fruits already taken... also very dubious.
Ar-Curunir 2 days ago||
Some of the problems solved here, at least in CS, have been open for decades, and have been worked on by very smart leading researchers in the field, including Turing Award winners.

Like, these would be best-paper awards at many top CS conferences.

asdfologist 2 days ago|||
Unlike math, software is constrained by the physical world.
cvak 3 hours ago||
In what sense?
skybrian 2 hours ago|||
This will depend on the problem; I expect big algorithmic performance improvements in AI since the algorithms are still new, inefficent, and constantly being improved. But maybe not for sorting, fast fourier transforms, or other well-studied basic algorithms?
pavpanchekha 2 hours ago||
A lot of algorithmic improvement in AI is ultimately bottlenecked by compute. It is very easy to come up with ideas that could improve models! But to prove that they do, especially at scale, is expensive and takes a long time.
slashdave 2 days ago|||
> we are also seeing incredible advances in software performance

Incredible?

> open ai announced like 15% improvement by fixing gpu kernel issue

That is... ordinary software optimization.

blovescoffee 2 days ago|||
a 15% improvement at a trillion dollar scale company is massive
enraged_camel 3 hours ago|||
There's nothing ordinary about downloading a new GPU driver and having performance go up by 15%.
obidan 2 hours ago||
This is untrue. It is very ordinary. How much do you know about GPU drivers that you state this so assuredly? Drivers are software. Software can be improved. Do you believe there are no prior examples of GPU drivers being improved such that particular compute patterns go up in performance by more than 15%? This driver improved Total War performance by 71% https://www.nvidia.com/download/driverResults.aspx/74714/en-... Also note you can go ahead and improve any open source driver right now, most likely. Compile it for your specific card and remove all other architecture specific if-cases and you can get an improvement.

Edit: also here’s a opencl 30% compute perf increase documented here : https://m.hexus.net/tech/news/graphics/74425-haswell-systems... that i just googled for

dominotw 2 days ago||
> novel advances in math

> we are seeing frontier level math breakthroughs (ie performance that would put it in the top 100 or 1000 mathematicians in the world if it were a human, meaning top .00001% or 800/8B)

i think you have misunderstanding of what mathematicians do

DaiPlusPlus 2 days ago||
> i think you have misunderstanding of what mathematicians do

They get to make cool 3D plot visualizations of functions so obscure to me that they’re named after someone who is still alive - and/or get to work on cryptography for the NSA - I think?

gpm 1 day ago||
Henry Yuen's (whose work problem 6 builds on) comments on this are worth reading IMO: https://bsky.app/profile/henryyuen.bsky.social/post/3ms2jpch...
jhrmnn 3 minutes ago||
This starts to feel like chess engines. It’s obvious their play is superior but it’s impossible for humans to understand the moves.
an0malous 2 hours ago||
It sounds like he hasn't verified the results of a problem that he has personally worked on, so how many of these problems have actually been verified?
margorczynski 48 minutes ago|||
From what I understand all of them have Lean proofs/certificates thus are basically 100% proven without a doubt.
samrus 5 minutes ago|||
We recently saw that lean itself isnt proven correct. Its not likely but i wouldnt call it verified if its only verified in lean

https://x.com/gro_tsen/status/2082483878480977959

voxl 17 minutes ago|||
Incorrect. The statement in Lean can itself be wrong. Moreover, they could be exploiting a kernel bug in Lean, of which we had one published literally a week ago.
gpm 1 hour ago|||
I mean, they're verified in the sense that the lean proof checks out... and presumably OpenAI read them.
doctorwho42 55 seconds ago|||
Or they made another LLM 'read' them?

> You are an expert in the field of mathematics, with decades of experience. You are a reviewer of proofs, etc etc.etc.

areoform 46 minutes ago|
Looking at this thread, I can see that a lot of technical people have ambivalent to negative feelings towards AI, but with each new generation, I become more and more convinced that they're missing out on something interesting.

It is indeed true that all models are, at their core, predictors of what occurs next in a sequence. But I think it's worth exploring the implication of what that means. Because when fed tiny pieces of information for a few tasks at a small scale, this results in something that sorta, kinda works. Or, works surprisingly well.

But when scaled... When the amount of information starts approaching the sum of all human knowledge, the tasks start approaching all useful applications of that human knowledge, and the fidelity of the predictor approaches incomprehensible sizes, the starts encodes / becomes (I'd argue it becomes) something that can model all human knowledge.

It feels wrong to say that, but let me explain, what is the best way to predict the behavior of a ball constrained in two directions that bounces with initial vertical velocity v(y) (y is up / down axis) and horizontal velocity v(x) (x is side by side in 1d) ?

If we purely look at it via a graph, it's by modelling the function of acceleration under earth's gravity.

If only a few points are given to you for this and you can't make something really sophisticated, then you'll make something that's rough that kinda sorta works and then call it a day.

But... if the number of points keeps increasing in number, precision and accuracy as well as the number of examples (assumed that data about air pressure, velocity and all other factors is included alongside these points), the fidelity with which you can replay / tweak the function keeps improving, and the number of times you can iterate keeps increasing, you'll eventually create a function that models that process so well that it intrinsically contains a good enough model of the deformation of the ball (provided the dataset contains information about elasticity of the ball's material, its dimensions and mass etc..), the nearly negligible (under normal conditions) effects of the ambient environment (provided there's diversity in the number of environments supplied), the oblateness of the Earth and minute changes in the gravitational field (the length of a seconds pendulum varies depending on where the experiment happens. It's presumed that all of the prior set of experiments were repeated across the Earth and the subtle, but real deviations were faithfully recorded)... and so much more.

A machine trained on the above with a large number of parameters, measures to prevent "laziness" and enough reps for high fidelity across a large enough dataset would start to approach a simulation of the ball falling. Because to predict what happens next in the sequence, you must model what's occurring in the sequence.

Now imagine doing that for other tangible and intangible things in this world. For all of human knowledge across all fields of endeavor. All experiences. No matter how noble, ignoble, notable or ignorable. But putting all of it into the soup that's this machine. Then at larger and larger scales, you eventually start encountering "good enough" models (in modelling the falling ball sense) for even the most hard to quantify / qualify things like grief and joy. At some point, by simply trying to predict what it has been taught ought to be the next part of the sequence in say... human interaction, it starts to make a model of something that hews ever closer to a full fidelity theory of mind.

Is there evidence for this? Kind of, yes. There are early indications that as machines are trained for an ever larger number of tasks at larger and larger scales, their internal representations converge. It's called the Platonic Representation Hypothesis. Overview and paper here, https://phillipi.github.io/prh/

It is my opinion that these machines are displaying a new form of intelligence that human beings haven't quite encountered before. They are the sum of all human knowledge made manifest and given voice by processes that nudge (bit-by-bit) what kind of step it ought to predict for the next part of whatever sequence it displays.

In my mind this means that, of course, these models can create new knowledge. This strains the analogy, but with the sum of all human mathematics within them, they can "reason" via the act of predicting what ought to come next.

Of course, these machines are "surprisingly" good at a lot of things the larger they get, because what the labs have created here is a rough version of humanity's collective knowledge given form and the ability to say hello.

I suspect that the current generation isn't close to the "true frontier" of what these machines could be. They are nowhere close to the sum of all human knowledge and endeavor. They are quite a way there, but they haven't yet achieved true completeness for domains where the data isn't so public.

I think it's the most exciting scientific and technological breakthrough of my lifetime. And I can't wait for us to get close to the true frontier of all domains.

More comments...