Top
Best
New

Posted by dcastm 3 hours ago

Are AI Labs Pelicanmaxxing?(dylancastillo.co)
208 points | 87 comments
simonw 1 hour ago|
This is fantastic

I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.

Catching a lab cheating specifically on my one dumb benchmark would be really funny.

Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.

His conclusion:

> Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.

lukev 1 hour ago||
What if they’re not pelicanmaxxing, but svgmaxxxing in general?

Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge.

Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

kaliqt 1 minute ago|||
Funnily enough, not that niche, because I have tried many times to do it as part of a wider project.
qq66 45 minutes ago||||
But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.
Balgair 7 minutes ago|||
https://www.youtube.com/watch?v=jgYYOUC10aM

reminds me of this Key and Peele skit

ryukoposting 37 minutes ago||||
How useful actually is this? It generates SVGs of pelicans on bicycles, sure, and some of them are (almost) spatially correct. But, none of them look good.

AI image generation suffers from this more generally. You can generate pictures of pelicans, sure. Newer models clearly generate images with more pelican-ness than before. But all of it is still uglier than sin. Drawing things accurately is one thing, making results that someone might actually want to use (without embarrassing themselves) is something else.

zahlman 21 minutes ago|||
Even if the models don't break through any particular "uglier than sin" barrier, with a bit more work, presumably the SVGs could become importable into an editor that would let a human apply taste and discretion. Seems to me like a heck of a head-start.

As for conventional diffusion-model stuff, I happen to think there are some pieces of AI art that still look really good even knowing they're AI.

amarant 16 minutes ago||||
Vibecoding a SVG based metroidvania as we speak! This is gonna be lit!
fiddlerwoaroof 29 minutes ago|||
AI can generate a fairly satisfactory SVG for a favicon now (programmer art quality at least).
sysguest 31 minutes ago|||
well that holds IF svgmaxxing is 100% "code-writing-maxxing"

...which.. hmm I dunno if they are same or not

tsimionescu 6 minutes ago|||
No, the point is that a general-ish ability to draw good SVGs is a useful ability in itself. People need SVGs for all sorts of purposes, and if AI can generate one for them, that's mostly useful (discussions about art and employment etc notwithstanding).

That said, I think this would correlate relatively little with general programming ability. They're not unrelated, of course, but being able to generate code that paints an accurate + esthetically pleasing image is quite different from generating code that achieves a non-spatial goal.

wasabi991011 27 minutes ago|||
I don't see why that's true. LLMs don't have to only be good at code-writing.
netsec_burn 1 hour ago||||
Addressed in the article, in case you're curious.
lukev 49 minutes ago||
Well, it’s mentioned as a limitation of the analysis, very much not ruled out (or in.)

That simonw is causing labs to do extra fine-tuning runs for this seems highly probable :)

charcircuit 1 hour ago|||
I agree, other formats, both textual and binary should be tested.
eob 22 minutes ago|||
Simon I hope from this day hence, your bio always includes:

"Simon Willison, among other things, is an advocate for the inclusion of pelican geometry in LLM training datasets."

docheinestages 38 minutes ago|||
I think a more fundamental test is SVG art creation in general. Perhaps a pipeline to take any image, caption it, ask the LLM for an SVG, rasterize to an image, and finally either use a deterministic visual similarity check or ask another LLM to be the judge and score how close the SVG is to the original image.
zahlman 19 minutes ago||
Fidelity to the original is definitely not how humans would measure "art" in this context.
docheinestages 8 minutes ago||
True, maybe we can call the generated SVG something else than art.
gopalv 1 hour ago|||
> Catching a lab cheating specifically on my one dumb benchmark would be really funny.

Similar thing happened when TPC came up with SQL benchmarks.

If you're not good at TPC, your engineering team is no good.

If you're good at TPC, then (as a customer) we will actually include you in a bake-off benchmark for our specific problem.

Winning on it is the price of admittance into the game, especially in a crowded market.

But how narrowly you benchmarket matters, you can't just hard-code that specific scenario & not fix anything adjacent while you're at it.

For example when it comes to GPUs, the "Quack3" (sic) benchmark on ATI cards comes to mind.

gilleain 1 hour ago|||
Perhaps also vary the bird? Wikipedia tells me pelicans are in the order _Pelecaniformes_ so shoebills or herons might do.
cyberax 32 minutes ago|||
> I've been casually spot-checking other animals in other vehicles

Snakes on a plane, weasels on a diesel, spiders on a glider, baboons on a balloon, goats on a boat.

mattertoast 1 hour ago||
[dead]
mauvehaus 1 hour ago||
> All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.

> However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest

Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any representation of a bicycle that shows the drivetrain you're going to show the right side of it if you want to do so without the frame occluding it. It's an excellent bet that their training data reflects this.

Citation: https://www.rei.com/c/bikes

Edited to add:

As near as I can tell, all of the bicycles are shown facing right, regardless of the direction the animal is facing (GPT 5.6-Terra, Sample 1/3). Also, in every case where the rider has legs (i.e. not the whale) both of the rider's legs are on the right side of the bicycle. This suggests a pretty serious lack of actual understanding of how a bicycle works.

stusmall 1 hour ago||
I'm glad someone ran the numbers on this. Every single Simon Willison post of an SVG is followed with someone dismissing it saying "I'm sure they train on it by now." This is despite a good blog post with sound logic on how easy that is to catch. [1] Glad to see someone took the time for a quantitative analysis of dumb little animals riding dumb little bikes.

1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...

elicash 15 minutes ago||
There's that version of the argument version, but there's also the softer version: that there used to be no training material of illustrated pelicans on bicycles, but now you have actual artistically talented individuals drawing it and that could improve the performance even though the AI labs are sucking it up no differently than everything else.

This post proves that hasn't happened yet, either. Although maybe the bad results posted online are being trained on and that explains the UNDER performance.

unholiness 58 minutes ago||
I don't think this small amount generalization to other animals and vehicles is strong evidence they haven't trained on this, either directly or more generally.
Wowfunhappy 2 hours ago||
> The more plausible story is SVGmaxxing

Exactly--and you have to ask yourself at this point what "maxxing" really means, since "get better at drawing SVGs" is a useful skill.

beering 1 hour ago|
Really awful how the AI labs are skillmaxxing /s

Pelicans aside, we need to remember that benchmarks are the only good quantitative way we have of comparing models. If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark! It is valuable work and very appreciated.

dbt00 44 minutes ago|||
It's a problem because of Goodhart's law.

If you train towards the test, you aren't necessarily improving overall fitness, but you are destroying the value of that test over time because you're decreasing its correlation with overall fitness.

Wowfunhappy 1 hour ago|||
> If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark!

I don't think I'd go that far!

When someone says a model has been benchmaxxed, what they really mean is that it performs better in benchmarks compared to their real world experience. That's a real thing, I've certainly experienced it with some models.

...my take is that some things in life just resist quantitative measurements. Who is the best job candidate? What is the best programming language? Add AI models to the pile.

bnfcl 1 hour ago||
This is funny, I actually did a similar experiment just yesterday.

Looking for evidence of the same, but with another twist: checking if the models would choose to create a pelican on a bicycle, if no specific bird or method of transportation was specified.

My version of it: https://www.modelbias.ai/pelican-on-a-bicycle-test

zahlman 4 minutes ago||
Simple as they are, there are some really aesthetically pleasing penguins on skateboards in there, including from less capable models. (In fact, I would say the Opus series got progressively worse at it over time.)
wasabi991011 20 minutes ago||
I find your analysis much more convincing than TFA, since it doesn't require a subjective evaluation and is more robust to animal/transport complexity.
bnfcl 4 minutes ago||
Thanks! Because I think that models are becoming better at creating SVGs in general. If you look at Claude Fable 5 and Kimi K3 for example.

In my tests it did create bicycles the most, but this is just a general bias I believe, as tested here: https://www.modelbias.ai/prompt/transport

dllu 1 hour ago||
I feel like getting LLMs to spit out an SVG is akin to getting a human artist to draw something by just reciting a list of coordinates. It's insanely hard and unnatural.

Image generation models nowadays can easily generate a photorealistic pelican riding a bicycle, where the bicycle has perfect structure. But it is, of course, only a raster image.

It seems that we're missing a kind of step to decompose an image into a list of instructions (say, SVG paths, or even brush strokes with a real brush) to reproduce it properly. Doing so would probably need a true understanding of the structure of the scene, which is something that AI still struggles with to this day.

cherioo 7 minutes ago||
I don’t quite agree. Good human artist can visualize in their mind how to draw a picture, i think. Which i think is no different than LLM doing SVG drawing in their “head”. Anthropic’s recent post call this head-space “workspace”.

It just might feel foreign to human who does not have a SVG trained head-space.

staticshock 1 hour ago|||
The pelican on a bicycle test is specifically about generating an SVG, fyi, not a raster.
dllu 1 hour ago||
I know. I'm just thinking about how to make AI create SVGs better... in theory, a sufficiently smart AI could "generate an image in its head", think about it, and then output the SVG paths to produce said image. Intuitively that would be somewhat closer to how human artists convert artistic visions into a sequence of arm movements while holding a brush (obviously, humans don't hold a fully formed, photorealistic image in the head while drawing, but rather vague concepts, but still).
0x000xca0xfe 1 hour ago||
Image models that support text output like Image2, or general text models that can read images like Claude can vectorize raster images. But they aren't very good at it, doing it manually in Inkscape still produces better quality even when done by non-artists.
munk-a 45 minutes ago||
It's a method to grade LLM output - as such it's something that will receive focus in correcting for. As soon as people who have a say in where funding is going noticed it as a metric the labs started caring about their performance in it. In the best case the labs are focusing on improving SVG capabilities in general and optimizing Pelican production as part of that initiative - but now that it's a known measure it is no longer reliable.
ertgbnm 1 hour ago||
I've had the feeling that labs aren't pelicanmaxxing specifically but that they do have some sort of RL environment for SVGs that they are letting the AIs overcook in. Specifically I'm thinking of the gemini 3.1 pro annoucnement that seemed to have a huge leap in animated SVG performance but not much else impressive about it.

So they aren't pelicanmaxxing but they are benchmaxxing in a way. The benefit of the pelican was originally that uplift on the pelican signaled an overall uplift on intelligence. I don't believe that is the case anymore and it is just another jagged edge of model intelligence.

nostrademons 30 minutes ago||
It's really refreshing to see someone publish a null result.
simonw 1 hour ago|
Underlying data is available on GitHub: https://github.com/dylanjcastillo/blog/tree/main/_extras/pel...
More comments...