Top
Best
New

Posted by milkshakes 10 hours ago

Ten advances in mathematics and theoretical computer science(openai.com)
440 points | 720 commentspage 7
drcongo 10 hours ago|
This thread has an absolutely wild points to comments ratio.
big_toast 9 hours ago|
tomhow explains they gave the story another shot here: https://news.ycombinator.com/item?id=49158443
Kelteseth 10 hours ago||
What's up with the upvote/comments ratio 8 to 337 on this post? Are the comments already also ai advanced? (/s?)
luciana1u 2 days ago||
[flagged]
baq 2 days ago|
I asked ChatGPT and it told me these aren’t not important /s
utopiah 2 days ago||
[flagged]
utopiah 2 days ago|
To clarify a bit due to the downvotes : this is not a research paper from a startup or a public frontier lab, it is just PR from a corporation, thus yes an advertisement. Downvote all you like it's still of no value.
Windchaser 8 hours ago||
> this is not a research paper from a startup or a public frontier lab, it is just PR from a corporation

If they're publishing the solutions to these 10 problems, then this is, essentially, the announcement of 10 research papers.

Yes, it's partly for reputation (as are many research papers), but that doesn't mean it's of no value. The best way to advertise is to show that you're providing value.

sashank_1509 2 days ago||
[flagged]
matteoraso 2 days ago||
It's simple economics. Building a robot to do your chores is expensive and only a small minority of people value their time enough to buy one. Meanwhile, SWEs are expensive and GPUs are (comparatively) cheap.
unknownian 2 days ago||
You shouldn't be getting downvoted for something that a majority of pure math and art enthusiasts believe to be true. The truth is many of these entrepreneurs and VCs are obsessed with AI not for money or human progress, but because it makes them feel closer to being a "god" rather than a mere mortal. Much of it (especially AI art) is out of spite for human creativity, which is done by mortals with limitations.
eadwu 2 days ago||
Taking the stance of moral superiority is kind of funny. And pure math and art enthusiasts don't think they are closer to being a "god" from understanding/"discovering" math?

Stop coping and deluding yourself mate.

To begin with, whether AI is the one doing the discovering or not makes no difference. Any "pure math" person would aim to understand regardless - and would be quite glad that they have a longer paved path.

Any mathematician in academic or industry is more than likely not a "pure math" person (tainted by capitalism).

AlexeyBelov 21 hours ago|||
> Stop coping

Isn't coping a good and useful mechanism?

xanderlewis 5 hours ago||
Exactly what I think every time someone uses the word 'cope' these days. It's like some kind of virus.
unknownian 2 days ago|||
Lmao what a ridiculous response. Yes, some mathematicians and artists are in it to feel smart. But the vast majority also just enjoy the process. Having a computer do all the work for you and just typing prompts in ruins that completely. As Ronny Chieng said in his Harvard speech, the journey is the point.

>Any mathematician in academic or industry is more than likely not a "pure math" person (tainted by capitalism)

Ignoring that I meant pure as in non applied math, let's just make it clear: you agree that mathematicians who are against capitalism encroaching on this process should be allowed to dislike it without criticism of being pretentious?

xyzsparetimexyz 2 days ago||
Any implication of any of these findings? They seem like unimportant nerd snipes to me. If you want to do something actually relevant, get chatgpt to write a simulation of graphene nanotube construction and figure out how to do it at scale.
utopiah 2 days ago|
Very marketable nerd snipes indeed.
foobar10000 2 days ago||
One - and I do not mean to be snarky - you can literally ask Gpt 5.6 Sol this - and if you want to see cool stuff - Fable running in their app (not website) has a view thinking button that is actually a good way to explore the adjacent fields, etc.

The non-sofic group one is definitely a big deal - would have been a Fields medal if discovered by a human.

zkmon 2 days ago||
> claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work.

AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.

Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.

raincole 2 days ago||
Sorry, OpenAI's take is correct here. If you're not convinced, here is how they prompted LLM: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98... [0]

A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.

[0]: Not one of the proofs in the linked article, but from OpenAI too.

ben_w 2 days ago|||
> A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.

I think you're over-estimating what a smarter highschooler could write.

A "finite loopless undirected multigraph" could have been explained to me at that age if we'd taken Discrete rather than Mechanics and Pure (and one module of Stats) in my two A-levels* in maths and further maths; but from what I saw of the Discrete module, neither:

  Every finite loopless multigraph with no bridge possesses a cycle double cover, without additional assumptions such as cubicity, planarity, connectivity, or higher edge-connectivity.
nor:

  repeated-edge closed trails masquerading as cycles
would have been something we'd have learned. But more importantly, we absolutely didn't have a feel for how much effort one needs to put into making sure the proof is right, so if one of us had been hypothetically asked to write a prompt it would've been no more than half that length, and missed most of the bullet points.

* For those not from the UK: A-levels are between secondary school and university, when aged 16-18. Functionally they are university entrance qualifications: https://en.wikipedia.org/wiki/A-level_(United_Kingdom)

yaqubroli 2 days ago||||
The human provides the intention and the ability to appreciate the output. Tools do “heavy lifting” all the time, but we still primarily credit the humans who use them precisely because they made the choice to use them.

Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.

esikich 2 days ago|||
What gives the intention and ability to the human?
oklahomasports 2 days ago|||
Are you playing dumb? Using power tools to build furniture is very different than using an ai robot to carve a statue or whatever.
ipnon 2 days ago||||
But why can’t we prompt the LLM “just do math research”? This is what I don’t understand.
ascots 2 days ago|||
100% agree. If the models are so capable that they're advancing math, it doesn't seem like a stretch to expect they should be able to determine with "doing math research" entails and the best way to use their capabilities towards that end. Why do we need to hand hold the models by telling them to do parallel research, keep threads independent, etc.
raincole 2 days ago|||
If there aren't thousands of TPUs doing that [0] right now I'd be quite surprised.

[0]: e.g. "go through wikipedia's unsolved math problem list and solve them".

mathisfun123 2 days ago||||
I don't disagree with you but there's no need for exaggeration; ain't no high school student writing this:

> In particular, proofs for special graph classes, constructions of cycle covers with some edges covered other than twice, bounded-length or prescribed-cycle variants, reductions to another unproved conjecture, computational verification through any fixed graph size, and candidate counterexamples without a complete nonexistence certificate are insufficient.

which is infact a very important part of the prompt.

don_esteban 2 days ago||
the fact that such things have to be explicitly in the prompt points to the fact that the underlying system is still far from where it needs to be (basically, lacks basic understanding what a proof is)
skinner_ 2 days ago|||
No, that's not what this is. This is a warning to the LLM that coming back with partial results is not good enough.

Take a grad student with a perfectly good understanding of what a proof is. Their supervisor gives them a major problem to work on. Almost always, the problem is too hard, the student comes back with partial results, and student and the supervisor iterate from there. Now imagine that they have an unusually cruel and unreasonable advisor who tells them, do not dare to talk to me until you've fully solved the problem. This paragraph is exactly that. It's there exactly because the underlying system is smart enough to know that real mathematicians do not work like that.

don_esteban 1 day ago||
If that was the case, that elaborate listing of all things that might look like a proof to a naive student, but are actually not proofs (and not even just partial results, but fundamental misunderstandings of what constitutes a proof) could have been easily and equivalently replaced by 'I am interested only in a full proof, don't bother me with partial results'. Yet, they were not.

To a real mathematician you would not have to list those explicitly, he/she would have understood that implicitly from 'give me a full proof'. That listing makes sense to say only to somebody who pretends to be a mathematician, but has not true understanding of how the math works. The models are getting better and better in this pretension, but prompts like that reveal that it is still just a pretension, not a true understanding.

famouswaffles 2 days ago||||
They don't have to be. At this point, we have multiple results from 3rd parties where the prompts are very basic.

To name a few:

- https://xcancel.com/DmitryRybin1/status/2079904005652893709

- https://archive.ph/2w4fi (https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...)

mathisfun123 2 days ago|||
you cannot be serious
zkmon 2 days ago|||
When you use a crane to do the "heavy lifting" for construction work, do you give full credit to the cranes?
raincole 2 days ago||
Read the prompts in the PDF I link and see if your analogy makes sense in this context :)
zkmon 2 days ago||
Prompt quality should not matter. If a high-schooler operates the crane to lift a ton of weight 10 floors high, should the credit entirely go to the crane?
Anon1096 2 days ago||
When I type 56789*23456 into my calculator and get the result I don't claim to have solved the problem, the calculator did it.
par1970 1 day ago||
qed
esikich 2 days ago|||
Your brain also is physical. Electrochemical gradients flow between physical molecular constructs. Isn't it just chemistry? Do you attribute it to physics or some whole-is-greater-than-the-parts idea?
NitpickLawyer 2 days ago|||
A better analogy would be a manufactured object, say 3d printed for simplicity. The 3d printer is given an input, and an object manifests itself after some time. We say that the creator of the object is the person turning on the machine, sending the data, and collecting the object. Not the machine itself.
cure_42 2 days ago||
I'd say the creator is the one who created the 3d model, not the one who pushed the print button.
samatman 5 hours ago|||
True story: I have a moisture issue in my furnace, such that it needs vacuuming out. This involved detaching a length of tubing, but that puts stress on said tubing, sometimes knocks other things out of alignment, and involves completing the seal between the wetvac and the tubing with my hand.

I also have a 3D printer. I also have a ChatGPT subscription, and some OpenSCAD chops. I came up with a part which would go into the top of the down tube to the drainage pump, and mostly-seal the down tube itself, with an opening on the side to vacuum out the moisture. This was purely prooompted, I took some measurements, printed bits of the part, refined the shape, and you know what?

It works! I can stick it down there, turn on the (very loud) wet vac, and go upstairs. On a 1.5Ah battery it sucks for a bit less than ten minutes, which turns out to be plenty of time.

So: who made that?

Don't care. I'm waking up warm at night.

Also: me, obviously. ChatGPT doesn't have a fucking furnace.

dgellow 2 days ago||||
I would say „I made this gadget with my 3d printer, but the designer is someone else (I found the model online)“. The intent, the drive, the action comes from the human
traes 2 days ago||
"I made this proof myself, but the designer is someone else" is an extremely unconvincing claim to ownership.
dgellow 2 days ago||
Almost as if a proof isn’t the same as a 3d print. It’s just not a good analogy
NitpickLawyer 2 days ago|||
(let's assume that)My 3dprinter is special. It has a bunch of values + an algorithm (i.e. a neural network) that takes input as tokens and outputs a printed object.
ben_w 2 days ago|||
> Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.

When the tool is a 3D printer, or any CNC system really, you bet I attribute a build to it.

I could also attribute the operator; there is no contradiction, it's a free choice, just like saying "I am in Berlin" does not contradict "I am in Germany".

naasking 2 days ago||
> AI has no self-awareness

What is your mechanistic model of self awareness that yields this conclusion?

> It's a tool

Does your model suggest that tools can't have self awareness?

perching_aix 2 days ago|||
Dunno about the parent commenter, but I personally interpret the concept as having a hidden representation of self that is continually tended to, and influences future choices. This implies statefulness, which models are intentionally not at inference time (*).

(*) Even if we hack around this and just do the usual trick of simply laundering statefulness to a higher level, in this case the context window being fed in, I fail to identify (**) a representation of its own state in these bodies of text that it'd be meticulously maintaining. I further fail to identify how it could be hidden or maintained, considering I control like half of it. The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either. Difference being, I actually have one (I'm "on-line").

You'll sometimes catch models mixing up who's who and how many who-s there even are for example.

(**) I did wish for something hidden though, so maybe it's just concealed? The same way people can encode a lot more of their emotional and mental state than normal into text if they read and write a lot of it, I'm aware of research that suggested the same for LLMs, albeit I cannot cite it. Maybe those phrasing signatures are just alien to me and will never pop out. Either way, I'd expect researchers to stumble upon this during interpretability studies, and either they haven't, they have but it wasn't popsci adopted, or they're keeping awfully tight lipped about it. If you know of anything like this, your turn now, would be happy to learn.

I do wonder how reasonable it is to expect e.g. a single maintained identity though. Maybe it isn't?

(*) Another way to hack around this of course is to just precompute some internal "self-awareness states" and hop around between them. Probably the closest to what the models are actually "doing".

ben_w 2 days ago|||
Before reading, know that I am uncertain in either direction.

> a hidden representation of self that is continually tended to

This sounds like a personality? They act like they have one of those. It may be an illusion, and even if it isn't an illusion it is unlikely to be anything like the source (us), but they act like it.

> I further fail to identify how it could be hidden or maintained, considering I control like half of it.

Indeed you control everything about a local model, and much of the context of even a remote model. But the state of activations and circuits in SotA AI is hidden in similar ways to those of synapses in your head: difficult to decipher even with probes monitoring the signals directly, and often not emitted at the normal output.

> The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either. Difference being, I actually have one (I'm "on-line").

While we can be confident that LLMs make up personas etc., it is insufficient to go from "that doesn't necessarily represent the model's internal state" to "therefore it doesn't have one".

> You'll sometimes catch models mixing up who's who and how many who-s there even are for example.

I've, unfortunately, also experienced this with humans. Perhaps they were losing their self-awareness at the time? I do wonder if old-age dementia does that by the end, though the person in question didn't ever get diagnosed with that.

> If you know of anything like this, your turn now, would be happy to learn.

Do you mean like these, or something else?

• https://researchportal.hkust.edu.hk/en/publications/decoding...

• https://aclanthology.org/2026.eacl-long.165/

• https://transformer-circuits.pub/2026/emotions/index.html

perching_aix 2 days ago||
> This sounds like a personality?

Not quite what I meant, but it's also not entirely unrelated I guess? Personality to me is like a natural bias. It does also shift over time, and is also an internal bit of state. I guess in some respects it can also be self-referential, like personal convictions.

> Perhaps they were losing their self-awareness at the time?

I do think it is entirely possible for people's self-awareness to shift, yes. Or more precisely, I do model things that way.

> Do you mean like these, or something else?

They're adjacent, but I more meant something like these:

https://arxiv.org/abs/2410.03768

https://arxiv.org/abs/2310.18512

https://arxiv.org/abs/2605.26537

So basically, steganography. The difference is that these papers investigate from the perspective of separate LLM instances covertly exchanging information between each other. This is in contrast with the scenario I'm laying out, where an LLM's past state is exchanging information with its future state, continuously representing and modulating a concealed internal state of some sort. And then that state just so happening to be some sort of self-referential meta state.

And the best inkling I have towards this is basically: https://www.youtube.com/shorts/WP5_XJY_P0Q

But then I don't think there's enough covert channel bandwidth in the agent replies for anything interesting like this.

naasking 2 days ago|||
> Dunno about the parent commenter, but I personally interpret the concept as having a hidden representation of self that is continually tended to

I don't see why an LLM could not have a sense of identity or personality while it's evaluating a specific prompt, or even change self awareness while evaluating a prompt since many outputs model a back and forth conversation. My point is that without a mechanistic model of what "self awareness" means, we have no way of truly evaluating such questions, we're just hand waving vague intuitions about what it could mean.

perching_aix 2 days ago||
Sure, but then such a model is not going to make itself. People pitting their vague intuitions is how such models eventually form. I'd also push back regarding that my comment would have been handwavey or without mechanistic elements, even if it was on the whole informal.

This is kind of also the reason e.g. the HN site guidelines are worded the way they are. Regrettably, forums naturally yield themselves to tit for tat type exchanges, but there's really no reason one could not bounce such vague intuitions off of another. I do not have to be right or wrong, and you don't either. Admittedly difficult when its some intensely contentious topic.

If a mechanistic model existed, there would also be no reason to talk about this in the first place. There'd be nothing to discuss, you'd be simply told how a given model characterizes from this perspective on the model cards.

naasking 1 day ago||
Even mechanistic models generate interesting discussion. How many years have we discussed Turing machines and the lambda calculus? Almost a century of great work came out of those.

The reason I insist on mechanistic models is because the original post was making a definitive knowledge claim, and in my experience, the knowledge claim is unwarranted.

Delk 2 days ago||||
I honestly don't think a language model is enough for self-awareness, regardless of the exact model of awareness.

A language model (or an image model or whatever) cannot even be sentient, and I think sentience is a prerequisite for awareness.

Even if we express a lot of our subjective experience with words, the language is just a symbolic representation of those experiences. The qualia themselves, even those that are quite abstract, are rooted in our physical presence and evolution.

You can't have an understanding of what hunger or physical pain feel like if you have no need for food or a sensory capacity for feeling pain. You can't understand what loneliness or pride at an achievement mean if you don't have a neural network wired to value social connection or status. We value connection because we're social animals that have needed each other for survival.

Even the more abstract of our subjective experiences are in some way rooted in our physical evolution.

I see no reason to believe that a neural network built entirely based on the symbolic level of language could have the features needed for the subjective experience itself.

AI awareness might actually be more believable if that awareness manifested itself in an entirely different way than in humans. But if we assume awareness because outputs resemble what we consider meaningful as humans, yet the neural network has had no inputs or evolution that could form the actual basis of human-like experience, I think we're seeing something that isn't actually there.

naasking 2 days ago|||
> The qualia themselves, even those that are quite abstract, are rooted in our physical presence and evolution.

There is no objective evidence of qualia. All evidence of qualia are vocal or other expressions of belief in qualia. Perceptions clearly exist and are observable, subjective experience and qualia, not so much.

> I see no reason to believe that a neural network built entirely based on the symbolic level of language could have the features needed for the subjective experience itself.

If your objection is to models based on "symbolic level of language" which you think lack semantic understanding of, say, trees, you should ask yourself how our brain, based on physics which also lacks any semantic category for trees, can somehow develop a semantic understanding of trees. All of these appeals to differences with the brain never seem to acknowledge that fundamentally, the brain has the same explanatory gap with physics.

> But if we assume awareness because outputs resemble what we consider meaningful as humans, yet the neural network has had no inputs or evolution that could form the actual basis of human-like experience

This assumes a lot. It seems very possible to me that intelligence inherently develops a map of natural categories (natural kinds), and language naturally develops around such categorical understanding. Semantics are then fundamentally the network of associations between categories, eg. there is no fundamental difference between symbols and semantics, and the latter cam be inferred from the former, and that's exactly what LLMs do, and why the semantic maps between different languages are so similar and how they can translate between languages.

woeirua 2 days ago|||
So… your model is 100% vibes based. Got it.
Delk 2 days ago||
I wasn't trying to give a model. The point was that I don't think it's necessary to give one.

You didn't address any of what I wrote, let alone provide any counterarguments. Which part of what I wrote do you think was wrong?

naasking 2 days ago||
Making definitive claims about whether LLMs do or do not have specific properties absolutely does require precise definitions of those properties that can be used to evaluate those questions. Merely hand waving that LLMs didn't undergo the same evolutionary process is not a definitive argument.

For example, the Turing machines and the lambda calculus don't look anything alike, but they are fundamentally interconvertible, and so in a real sense they are fundamentally equivalent. Without a model, all of your arguments are completely unconvincing for exactly the same reasons, eg. that there may exist many paths to fundamentally equivalent ends.

Delk 2 days ago||
I just don't think linguistic (or other symbolic) representations alone can contain the information, in any sense of the word, of what e.g. human subjective experiences actually are like. The concepts we express with language get their meaning from our physical reality, even if quite indirectly in case of some abstract concepts.

Hunger as a concept doesn't mean anything without the physical need. Politeness or bluntness, even in writing, don't mean anything without social dynamics. And we have social dynamics (and neural structures that directly process social cues and associated feelings) because we've evolved into social animals for whose survival that was important.

I see no reason to believe that a model trained only with symbolic representations, with no connection to the physical world phenomena that those symbols represent, could contain the subjective experience itself.

Neural network models may be able to derive novel (or at least novel-looking) output rather than just an obvious rehash of their input, but I don't think any set of bytes can fundamentally contain information that was never entered into it. (Even if e.g. a model produces previously unknown mathematical results, those results can in principle be derived from the information that they were trained with.)

I'm not saying that artificial neural networks couldn't, in principle, be aware. ANNs and biological neural nets may be equivalent in the sense that any information and processing structures represented by a biological one could in principle be represented by an artificial one. If that's the case, and awareness is purely a product of our neural systems as materialism would imply, it should be possible for an ANN to be aware, too.

But when the model has been trained with only language, and IMO the subjective experience can't be derived from the symbolic representation alone, I can't see how the model could include the actual subjective human experience.

An AI model could of course have an awareness and subjective experiences that are totally different than our human experience. But then the fact that it happens to produce output resembling what humans find meaningful shouldn't be considered indicative of such awareness.

This is of course more of a philosophical argument than a technical one, and I'm happy to hear counterarguments, but not on the level of off-hand dismissal.

naasking 1 day ago||
> I can't see how the model could include the actual subjective human experience.

People who say LLMs have subjective experience aren't saying they have human-type subjective experience. Nobody who sees an LLM express hunger when role playing as a hungry person thinks that the LLM is actually hungry.

I too can role play as a hungry person despite not being hungry, so there is no reason in either case to conclude that the words produced reflect genuine internal subjective states. The point is that such internal states may still exist.

overgard 8 hours ago||
Gary Marcus' has a good take on this:

https://garymarcus.substack.com/p/openais-amazing-but-vastly...

https://garymarcus.substack.com/p/two-critical-updates-re-as...

Not that there isn't something interesting in here, but lets be clear that we don't have enough information to evaluate this properly. And as always with these labs, BS takes a lot more energy to refute than it does to spread.

scarmig 8 hours ago||
It's worth reading Marcus' first line, for the naysayers and flaggers on this post:

> Astra, a new model that OpenAI is testing internally, is amazing. No denying that.

aaroninsf 8 hours ago|||
I find Marcus on this, something approaching sophistry and rhetorical showmanship in service of maintaining an ideological position, for reasons unrelated to the nominal intellectual clarity.

To sharpen that, I think he's (obviously) interested in maintaining his own brand as "thought leader" and this necessitates de rigeur defense of particular postures.

Sometimes this is easy because the facts warrant it; other times, a bit of rhetorical license is required to preserve nominal coherence and (at least, for the moment) hold certain lines.

This is one of the latter cases, and it's not subtle.

One of the celebrated properties of many intellectual advances or inventions in whatever domain is precisely that it appears obvious in hindsight. It is quite cynical to leverage consensus distrust of large AI players, warranted but also a popular social construction, to insinuate that these are not "real" advances or "real" hard problems, on the grounds they were in some sense cherry-picked.

Identifying the problems amenable to strategies on the table and intuitions (sic) about where bridges might be, is exactly the discerning work that is the core driver of almost all prior progress, but for celebrated accidents and flashes of insight. Anyone working in any challenging discipline knows that those are celebrated and told around campfires precisely because meaningful durable results arising like that is so uncommon.

These two articles make me think of nothing so much as my own durable reaction to the creeping goalposts of AI critics generally: that they often seem to me not unlike a water color cohort scoffing and jeering at the horse, because it got a D on its tensor calculus exam.

Marcus should be on guard against his own cynicism and take care that his assumptions do not prevent clear sight.

HardCodedBias 8 hours ago||
"Gary Marcus' has a good take "

I think that is an oxymoron.

neta1337 8 hours ago|||
How so? His predictions were accurate so far
energy123 8 hours ago||
No they were not. These were his 5 predictions in 2022:

""" 1. By 2029, AI will still be unable to watch a movie and accurately explain the characters, events, conflicts, and motivations.

2. By 2029, AI will still be unable to read a novel and reliably answer questions about its plot, characters, conflicts, and motivations beyond what is stated literally.

3. By 2029, AI will still be unable to work as a competent cook in an unfamiliar kitchen.

4. By 2029, AI will still be unable to reliably create more than 10,000 lines of bug-free code from natural-language instructions or interaction with a nontechnical user, excluding simple assembly of existing libraries.

5. By 2029, AI will still be unable to convert arbitrary mathematical proofs written in natural language into symbolic form suitable for formal verification. """

There's still 3 years to go and he's already wrong on 4 out of 5.

sweezyjeezy 8 hours ago|||
Well I don't typically side with GM, but playing devil's advocate:

1. still not wrong? Unless it's just feeding the audio or screenplay I don't think you can feed AI a full movie in a single context window yet?

2. Not sure, but can you prove this wrong? Can you feed a full, unseen new book and get that kind of answer?

3. Not wrong.

4. I think he'd probably pull you up on 'bug free' - I don't think that frontier models can reliably write 10k LOC without _any_ bugs typically (not that humans can do this either).

Philpax 6 hours ago|||
4. I think they can, especially if the problem statement is well-specified and, importantly, autonomously testable. Of course, specifying a problem that meets these requirements is non-trivial, but the claim requests _a_ counterexample :P
sweezyjeezy 6 hours ago||
The wording was 'reliably' though? I could just be splitting hairs on that one though to be honest.
lostmsu 6 hours ago|||
1 is wrong. If I tell Codex + GPT-5.6 to do it now, it will figure out how to do it. If it would need to extract audio and run a speech model on it, it will find one, set it up, and run without my help.
sweezyjeezy 6 hours ago||
I'm not buying this. GM clearly was trying to set a benchmark for video comprehension, not tool usage. Video comprehension is required for many 'AGI tasks', especially robotics to work in real time.

An LLM could theoretically try to earn some money and pay a human to do all 5 tasks but it's clearly not the spirit of the challenge.

an0malous 8 hours ago||||
> There's still 3 years to go and he's already wrong on 4 out of 5.

Have these been tested or are you just guessing?

ducktective 7 hours ago|||
Do LLMs generate deterministic or trustworthy answers?
overgard 8 hours ago|||
It can be annoying when someone you disagree with is frequently right!
maxprimes 9 hours ago||
I'm sure OpenAI is just interested in the greater good of mankind!
merelydev 7 hours ago|
Great stuff. Wonder how many of the ten problems where solved by independent mathematicians not linked to OpenAI
p1esk 7 hours ago|
Zero. These were open problems.