Top
Best
New

Posted by delichon 16 hours ago

Karpathy’s Pelican(twitter.com)
https://xcancel.com/karpathy/status/2083749667410727319
229 points | 179 comments
YmiYugy 32 minutes ago|
I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.
jmugan 2 hours ago||
A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
maxutility 2 hours ago||
Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting.

A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.

IsTom 53 minutes ago|||
Aren't they still bad at understanding how bicycle frame works? Especially the steering part?
gegtik 42 minutes ago|||
Maybe that makes them human..

https://www.booooooom.com/2016/05/09/bicycles-built-based-on...

zh3 42 minutes ago||||
Shhh...you'll alert the models :)

Totally agree though, anyone with a vague understanding of how bikes works ignores the pelican because they know the bike is unrideable in the first place.

edaemon 46 minutes ago|||
Yes, but I think the idea here is that most models produce very similar pelicans on bicycles, so a different test might be more useful in gauging the differences in models.
dofm 1 hour ago|||
But it's another benchmark on how good models are at generating intensely average, unwanted things with unthinking design. Just scaled up.
charcircuit 42 minutes ago||
Bad? It has a charming style. I would watch the whole book if it was made like this.
altmanaltman 7 minutes ago|||
Very few things are universally hated. One can love something truly that is hated by most. But it doesn't change the fact that it's still hated by most. An objective and a subjective opinion can exist at the same time on this.
trial3 30 minutes ago|||
yeah, definitely, in the same way that we all regularly go and look back fondly at our chatgpt ghiblified family photos
bredren 3 hours ago||
I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page.

That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right.

But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them.

My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc.

Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs

I can share some of the Apocalypto bit if anyone is interested.

garganzol 44 minutes ago||
The product you've built the animation for is only for macOS, which is a pity cause it has a way wider appeal than macOS market share is. You have low-hanging fruits right there, don't miss them.
trvz 31 minutes ago||
macOS covers 90% of people who would ever pay him.
garganzol 14 minutes ago||
This is a bizarre claim to make in this case, AI tech is used on all platforms universally.
thejazzman 5 minutes ago||
iOS users spend dramatically more on e-commerce. I worked in e-commerce. Maybe it’s changed in the last 5 years. But that’s where the ops sentiment comes from.
manofmanysmiles 2 hours ago||
I am interested, I'd love to see! I'm waiting for the day my dad's self published books become self produced movies!
HarHarVeryFunny 2 hours ago||
It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.

When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.

fasterik 1 hour ago||
Taking a single paragraph of literary text, which is abstract and ambiguous, and converting it into a 3D animation requires an enormous amount of implicit knowledge about spatial relationships, intuitive physics, everyday objects, and so forth. Not to mention the mathematics of 3D transformations and computer graphics more generally. Saying that it's indicative of no more than three.js coding ability is absurd.
beepbooptheory 34 minutes ago||
Why does it require knowledge about spatial relationships?
fasterik 5 minutes ago||
Because the degree of realism in a video is determined by implicit knowledge of concepts like "in front of", "behind", "next to", "inside of", "between", "occluded by", etc. as well as distances, angles, and the relative size of objects when viewed from a perspective.
onion2k 29 minutes ago|||
I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.

This weekend I've been converting a game from three.js to ogl.js in order to see if I can optimise the time-to-interactive loading time. I took the three.js driven page weight from about 600KB (500KB being three.js) to about 50KB, and reduced the loading time from multiple seconds on a 4G mobile connection to around 0.5s.

This has mostly been a combination of Opus 5 and Sonnet 5 in Claude Code. It very clearly has a good grasp of WebGl, and of what impacts page loading times and rendering speed. It was able to drive Claude Code's integrated browser to measure the impact of changes, and as I spiked out a test of ogl.js it could test the differences changes made.

It's not the best game (https://tinyslots.ooer.com) but that's on me. As an exercise in building 3D in a browser, and in page speed optimization, with Claude models I am really impressed.

lowbloodsugar 1 hour ago|||
Don’t worry. I’m sure they’re not training it to be good at things like writing database backends, financial services, logistics systems, user interfaces, or anything of economic value. As long as you’re not working on three.js specifically, I’m sure Anthropic isn’t making any progress you should be worried about.
abletonlive 59 minutes ago||
[flagged]
peterleiser 41 minutes ago|||
But LLM's passed the Turing test, so we're done, right? ;-) Seriously, though, I agree with you that people keep moving the ball.
Applejinx 30 minutes ago|||
Oh, they still are, it's just that it's a rare parrot that can reproduce 'all of the written corpus of three.js' or stochastically align all that with arbitrary and weird conditions for what to squawk out.

This is what you get. It's like the more general concept of starting with the written word (or 1000 words) and then replacing it with a picture. You've done something strikingly different, but is it serving the same function?

It's fascinating to see this stuff combine such disparate sources in unexpected ways. But it is parrot, just not in the way you're expecting.

try-working 16 minutes ago||
this is not a good benchmark for models, but it's great if you're optimizing for attention on twitter because video content and 3d animations perform best on social media.

a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.

qwertox 3 hours ago||
I'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be.

"Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0

baxtr 52 minutes ago||
It’s a fun idea to re-animate dead google products by feeding product videos to an AI
misiti3780 3 hours ago||
i remember this failing, but this product looks pretty useful in 2026
throwaway27448 3 hours ago|||
It looked pretty useful in 2009. Failing to push it was baffling then too.
QuantumNomad_ 3 hours ago|||
It was open sourced as Apache Wave when Google shut it down.

Years later Apache moved it to read only because of low community activity.

The archived git repo on GitHub remains available to clone and revive as a fork.

https://github.com/apache/incubator-retired-wave

throwaway27448 2 hours ago||
This doesn't explain the lack of marketing. Apache isn't exactly known for pushing tech. Google never tried to improve on gmail.
QuantumNomad_ 38 minutes ago||
Sorry if it wasn’t clear, I was adding this info for context and for anyone who hopefully feels inspired to pick up Wave and make something from it given that it was all open sourced and all. And I figured that in the chain after your comment about it having been useful looking all along was a natural place to add this additional info and link.
khazhoux 1 hour ago|||
This was the first big secret “you’re not allowed to know what this is or talk to anyone about it” project inside Google. They wanted to be left alone and especially to not have to integrate with the mail team. And everyone heard rumors of gigantic bonuses if they hit whatever milestones, which wasn’t a thing for other projects. All this made them isolated within the company, and when the project was an obvious flop, you didn’t see anyone rushing over to help them.
patwolf 2 hours ago||||
I used it to plan a group beach trip back when it came out. We had a single page shared with everyone going on the trip. The page had live shopping lists, maps, weather forecasts, and other snippets of useful information. Now in 2026 I still can't think of any single technology that provides the same utility. Although to be fair, this might be a case of rosy retrospection.
tomjakubowski 2 hours ago|||
Notion pages are pretty good for shared trip planning docs. Although maybe without so many live updating widgets.
tikhonj 3 hours ago||||
notion isn't too far off

wave failed for weird google organizational reasons far more than anything inherent to the product or tech

dgellow 2 hours ago||||
It was awesome when released. I used it a lot, the multiplayer experience was awesome, and the mix of document-forum-wiki is still something I miss
xnx 54 minutes ago|||
It was too far ahead of its time.
dundarious 2 hours ago||
I can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.
kiwibyproxy 1 hour ago|
which to be fair, happens just a few paragraphs later :) I first watched without sound and thought "oh that's the birthday speech disappearance"
Waterluvian 28 minutes ago||
Speaking of benchmarks has anyone given AIs Where’s Waldo pages and asked it to find Waldo?

I’ve been trying it on them all and can’t find one that does it consistently. The best will tell me they can’t. The worst confidently point out one of countless Waldo-likes.

hooloovoo_zoo 9 minutes ago||
I suspect LotR is a singularly unrepresentative choice here considering how much info exists about it.
fzeindl 3 hours ago|
Regarding the argument about LLMs having difficulties auditing their work:

I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.

techblueberry 15 minutes ago||
To a certain extent to what end though,

On the one hand yes, almost every task I work at now is one off one of scripts I throw away.

The question is - where does the software the spec or the “code”.

A really complex game will probably always be token heavy. At least for the next few years code is still not free.

But certain software is just iterative by design. If we mean we regenerate all the for loops of a game from scratch, sure but I think “code” Is really more spec then implementation, and we’ll want to continue building things through iteration.

And even on the for loop point - Do you really want to spend millions of tokens rewriting a game every time you need to make balance changes?

qrios 2 hours ago|||
I'm sure this will become the standard. And plastic is an excellent analogy. Maybe we can take it a step further and compare it to on-demand 3D printing.

Why would anyone still use off-the-shelf software when they can have a system that has access to all data, can transform it into any form, and can export it in any format?

After years of thinking that I needed to develop a decent movie management system for my own films or a columnar browser for large CSV files, Claude and Qwen each delivered exactly what I needed in just a day.

lowbloodsugar 1 hour ago||
This. This is why the burst of posts on HN of “I made this useful tool/crate/application” were just so sad. The old model was getting what the kids now call aura by developing useful open source products: products where it’s far easier for someone to consume the product than write it themselves. We had people posting things as if that model still existed. Dude, you wrote it with an LLM! Posting them (here) is not only pointless, it’s advertising that the person who wrote it isn’t smart enough to understand that I have an LLM too.
8n4vidtmkvmk 2 hours ago|||
Yes. It's fantastic for one-off tasks.
gisely 2 hours ago|||
Does this make sense with economics of software though? Throwaway products compete with more durable versions of the same product because there is a cost per unit produced that can be minimized by using cheaper materials or production processes that cut corners. With software there is no cost per unit. There might be a market for one-off software that serves a very specific purpose where throwaway software can compete with adapting more carefully engineered software to that purpose, but I am not convinced there is a lot value in this market.
skydhash 2 hours ago||
> Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price

Have they? Most of the world production is tied down to expensive factories and machines. Yes, we have more products, but that the result of the global trade, which is a very complex system.

> Produce it cheaply and if it breaks throws it away and reproduce it.

I don't know why everyone would ever wants this. It's been parroted since forever, but the true usefulness of software is to be able to build it once and runs it indefinitely. If some edge case occurs, I fix it. Which is way cheaper than rebuilding the whole thing. The goal is to have something like OpenBSD's ed[0] or dmesg[1], which you only touch every few years or so

[0] https://github.com/openbsd/src/commits/master/bin/ed

[1] https://github.com/openbsd/src/commits/master/sbin/dmesg/dme...

More comments...