Top
Best
New

Posted by dcastm 5 hours ago

Are AI Labs Pelicanmaxxing?(dylancastillo.co)
265 points | 110 commentspage 3
jonatron 3 hours ago|
OK, so we've done animal_vehicle, how about new SVG ideas each time? I just tried "make an SVG of a man sitting in a chair at a computer behind a desk" which gives more interesting results than the animalVehicle test.
ninju 3 hours ago|
There probably good set of images of that description already so it does exercise the inference capability of the model
RobRivera 1 hour ago||
Chasing metrics Chasing dragons

Tomato, tomato

Rooster61 3 hours ago||
I find it humorous that the animal + plane combo appears to be such an outlier. I assume this is due to the models assuming the user mean plain and misspelled it in the prompt.
ramses0 2 hours ago||
I think it's actually due to "pelican on a plane" isn't the same as "pelican on an airplane" (Sonnet5 @ Flamingo x Plane), some consistent and warranted semantic/linguistic confusion!
flsw 2 hours ago|||
I noticed this happens especially with herons. My guess is it's because the model links "heron" to Heron's formula and the Cartesian plane
NitpickLawyer 3 hours ago||
GLM has 2 combos of "on a plane" literally sitting inside a plane, with a window and a bit of wing showing. That's funny.
zahlman 2 hours ago||
... Is that not how it should be interpreted?
comrade1234 2 hours ago||
Hilarious. Could you imagine being a programmer at an AI company and this is your assigned task?
stri8ted 3 hours ago||
You seem to assume training on pelican would not result in improved performance on other similar tasks. Why?
altcognito 3 hours ago||
He didn't. That's why the article exists. You have to do the science to see if it does.

He was asking the question - do we see gains across other tasks? The underlying question was: Is the additional attention given to this specific task creating a false impression of progress?

HarHarVeryFunny 2 hours ago||
Either you've memorized the outline (or detailed component shapes) of, say, a horse, or you haven't. Memorizing the outline of a pelican isn't going to help you with the horse.

You could train a model to do something a bit different like a pencil sketch, or vector graphic sketch, of something given a photo of it, and expect that to be a generalized skill, but if you are asking the model to do it "from memory" then memorizing a pelican is no substitute for not having memorized a horse.

tomas789 4 hours ago||
Having an objective score is quite difficult. Maybe it would be better to do a pairwise comparison and calculate ELO?
NitpickLawyer 3 hours ago||
Just click through the models. At a glance (and highly subjective) I don't see anything jumping out as oom worse than anything else. I only noticed a model placing the animal inside a plane (with seat and small window) but other than that, they all seem similar inside each model to me.
javier123454321 3 hours ago||
If you want to, go ahead, but it seems to me the author already exceeded the energy expenditure that this question warranted.
andy99 4 hours ago||
If an AI researcher was going to pelicanmaxx, they would almost certainly apply the augmentations mentioned in the article during training, e.g. randomly selecting animals and conveyances. You’d want a model that generalizes well, just sfting in that specific prompt would be pretty bush league for a frontier lab.

I don’t have any reason to believe they are gaming the benchmark, just saying. I do find the idea of a data labeller having to generate thousands of svgs of different animals on different modes of transportation quite funny though.

cute_boi 3 hours ago|
At this point, I think there are so many pelican images in the pretraining data that drawing a pelican no longer makes sense as a model evaluation task.
dcchambers 4 hours ago||
It's incredible that each model has it's own style that remains relatively consistent throughout all of the different generated examples.
busymom0 2 hours ago||
How does attempt 2 by Llama 4 Maverick look like a bald eagle??
ck2 3 hours ago|
I am not sure if this is how it works but let's say there was a reddit thread talking about the pelican benchmark and in it someone posts mockup examples of what an ideal result would look like

aren't some LLM going to digest that thread at some point and indirectly learn from it?

basically my point is originally this was a good benchmark because it was an absurd never-seen-before thing, but now that it is in content, some models are going to get a benefit in education?

you'd need the "AI" equivalent of an old-school "google whack", something with no previous results

* https://en.wikipedia.org/wiki/Googlewhack

gbalduzzi 2 hours ago|
They ingest so much data that a couple of reddit threads do not move the needle.

It is the reinforcement learning that produces more tangible results with less data, but it is something that the AI labs specifically selects and it is not picked up unknowingly

More comments...