Top
Best
New

Posted by dcastm 4 hours ago

Are AI Labs Pelicanmaxxing?(dylancastillo.co)
208 points | 87 commentspage 2
apwheele 2 hours ago|
So this is not my experience at all for asking about simple SVG icons for web-pages. Here is one of the examples I have tried for in the past, make a simple cartoon SVG knife for a map icon for a crime map.

https://x.com/CrimeDecoder/status/2080008114615537766

Can see the images for ChatGPT/Claude (Sonnet 5), and Gemini are all quite bad.

Jagged edge of LLMs. How do you explain being able to generate very complicated shapes in the Pelican example but cannot make a much simpler icon without just alluding to it is in the training data?

Topfi 34 minutes ago||
Thanks for sharing another solid data point. I fear you won't get an answer from my experience [0]. Unfortunately, the blog post decided to forgo the very models that I found to be the worst offenders:

> Here is what Qwen3.6-35B-A3B via Openrouter provided for a sloth riding a skateboard: https://imgur.com/a/Dy8fvR5

> Like Grok 4 Fasts attempt at a mushroom in a rowboat, it is barely recognisable as anything despite both Qwen3.6-35B-A3B and Grok 4 Fast having no issue with more popular (i.e. benchmarked) examples. [...]

> And here is Opus 4.7 [which simonw claimed to provide a worse pelican vs Qwen], again via Openrouter: https://imgur.com/a/Qus1Enf

Anyone who hasn't witnessed such deltas either hasn't looked at enough examples, a sufficient variety of models, or both. And they are, unfortunately, not limited to "SVGMaxxing", but a wide range of evals.

[0] https://news.ycombinator.com/item?id=48951229

robocat 1 hour ago||
Does asking for a dagger help?
apwheele 1 hour ago||
If you look at the raster image ChatGPT generated, that is fine. It is just this example (and other simple SVG icons I have asked for) result in pretty bad SVGs. It just makes me highly suspicious that the LLMs are learning shape primitives and extrapolating to new shapes, vs just having a big dictionary of prior examples and stitching them together.
robocat 1 minute ago||
[delayed]
scosman 2 hours ago||
join me in building the ideal training set for pelicans riding bicycles: https://github.com/scosman/pelicans_riding_bicycles
BeetleB 2 hours ago||
Oh great! You've now made it a lot easier for LLMs to train on this dataset!

Your next iteration will need different animals and different transportation options. You'll run out after a few iterations.

anuramat 2 hours ago|
"benchmaxxing by generalizing" is not really benchmaxxing
comrade1234 1 hour ago||
Hilarious. Could you imagine being a programmer at an AI company and this is your assigned task?
jonatron 2 hours ago||
OK, so we've done animal_vehicle, how about new SVG ideas each time? I just tried "make an SVG of a man sitting in a chair at a computer behind a desk" which gives more interesting results than the animalVehicle test.
ninju 2 hours ago|
There probably good set of images of that description already so it does exercise the inference capability of the model
Rooster61 2 hours ago||
I find it humorous that the animal + plane combo appears to be such an outlier. I assume this is due to the models assuming the user mean plain and misspelled it in the prompt.
ramses0 1 hour ago||
I think it's actually due to "pelican on a plane" isn't the same as "pelican on an airplane" (Sonnet5 @ Flamingo x Plane), some consistent and warranted semantic/linguistic confusion!
flsw 1 hour ago|||
I noticed this happens especially with herons. My guess is it's because the model links "heron" to Heron's formula and the Cartesian plane
NitpickLawyer 2 hours ago||
GLM has 2 combos of "on a plane" literally sitting inside a plane, with a window and a bit of wing showing. That's funny.
zahlman 56 minutes ago||
... Is that not how it should be interpreted?
bluealienpie 1 hour ago||
AI rating AI? Am I missing something.
Nnnes 3 minutes ago||
I'm surprised (but not really) that you're the only comment I see even mentioning it. The ratings may be even lower quality than the SVGs.

Obviously they're all a bit cartoon-y, what else do you expect from SVGs. But I'm not convinced you could find a single human on Earth over the age of 4 who would seriously give the vehicle in GPT/1/whale/plane a 5/5.

Browse through the options a bit and the rest is not that much better. Grok/2/cat/plane, one of the more accurate planes, got a 2/5. For the most part, vehicles entirely missing do get a 1/5, except for whatever it is in Gemini/1/heron/plane scoring 4. Animals inside planes get completely random vehicle scores I guess.

The cats are all orange, except for a few of the skateboard cats that are black. I'm sure there's nothing to read into there...

Well, I've convinced myself that the next effective test of multimodal models will be whether their judgments of LLM-generated SVG airplanes are anywhere close to reasonable.

zahlman 57 minutes ago|||
That seems to be how we signal "objectivity" nowadays.
recursive 34 minutes ago||
Missing something? You're missing the boat! Have some AI-prepared koolaid before you get left behind.
stri8ted 2 hours ago||
You seem to assume training on pelican would not result in improved performance on other similar tasks. Why?
altcognito 2 hours ago||
He didn't. That's why the article exists. You have to do the science to see if it does.

He was asking the question - do we see gains across other tasks? The underlying question was: Is the additional attention given to this specific task creating a false impression of progress?

HarHarVeryFunny 1 hour ago||
Either you've memorized the outline (or detailed component shapes) of, say, a horse, or you haven't. Memorizing the outline of a pelican isn't going to help you with the horse.

You could train a model to do something a bit different like a pencil sketch, or vector graphic sketch, of something given a photo of it, and expect that to be a generalized skill, but if you are asking the model to do it "from memory" then memorizing a pelican is no substitute for not having memorized a horse.

johndough 3 hours ago||
Another point for consideration: Specialized SVG models create way better looking pelicans riding a bicycle. (E.g. Refract V4: https://jumpshare.com/s/8liB7Aiuoo3yucbWGXjZ mirror: https://postimg.cc/McV70p84 )
solarkraft 2 hours ago||
That’s an impressive image, but what a mistake it was to click the second link (on mobile without an ad blocker). I wouldn’t send it to anyone I respect ...
johndough 18 minutes ago||
Thanks for pointing that out. I haven't noticed any adds in years with Firefox and Ublock Origin extension.

I'll look for a better image host in the future. I guess the economic incentives makes them all turn bad after a while.

ACCount37 2 hours ago||
The name is "Recraft V4", and from looking it up: yeah, it sure seems like whatever black magic they use for SVG generation kicks ass.
johndough 22 minutes ago||
Oops, autocorrect. Sorry about that.
tomas789 3 hours ago|
Having an objective score is quite difficult. Maybe it would be better to do a pairwise comparison and calculate ELO?
NitpickLawyer 2 hours ago||
Just click through the models. At a glance (and highly subjective) I don't see anything jumping out as oom worse than anything else. I only noticed a model placing the animal inside a plane (with seat and small window) but other than that, they all seem similar inside each model to me.
javier123454321 2 hours ago||
If you want to, go ahead, but it seems to me the author already exceeded the energy expenditure that this question warranted.
More comments...