Top
Best
New

Posted by dcastm 6 hours ago

Are AI Labs Pelicanmaxxing?(dylancastillo.co)
318 points | 129 commentspage 4
dcchambers 5 hours ago|
It's incredible that each model has it's own style that remains relatively consistent throughout all of the different generated examples.
busymom0 3 hours ago||
How does attempt 2 by Llama 4 Maverick look like a bald eagle??
ck2 4 hours ago||
I am not sure if this is how it works but let's say there was a reddit thread talking about the pelican benchmark and in it someone posts mockup examples of what an ideal result would look like

aren't some LLM going to digest that thread at some point and indirectly learn from it?

basically my point is originally this was a good benchmark because it was an absurd never-seen-before thing, but now that it is in content, some models are going to get a benefit in education?

you'd need the "AI" equivalent of an old-school "google whack", something with no previous results

* https://en.wikipedia.org/wiki/Googlewhack

gbalduzzi 4 hours ago|
They ingest so much data that a couple of reddit threads do not move the needle.

It is the reinforcement learning that produces more tangible results with less data, but it is something that the AI labs specifically selects and it is not picked up unknowingly

j45 4 hours ago||
The models definitely seem to pay attention to the tests.

Since the tests can be generally gamed with directing descriptions at it non-deterministically, there's a greater chance the questions solution can be found.

Of course, hopefully the models are instead adding patterns and types of questions as well and it makes the models more capable, but it may be limited in how it transfers to other types of questions in breadth or depth.

cute_boi 5 hours ago||
https://playcode.io/blog/macbook-svg-benchmark

I think we should stop using pelican benchmark.

dllu 4 hours ago|
I disagree with this in the blog post:

> Every single one is a pelican, on a bicycle, first try. When every student gets an A, the exam has stopped grading.

Numerous pelicans and their bikes are clearly horribly malformed. In fact none of the bike frames are correct. Fable and Opus come close, but the top of the diamond is disconnected in Fable's case and the head tube is misaligned with the front fork in Opus's case.

And of course, as the parent post shows, labs don't actually seem to be training on the pelican bike case.

ErrantX 4 hours ago||
Agreed. And more; the Macbooks are pretty much the same - some are god approximations, some are terrible, all of them are recognisably a MacBook. And if you start using it they can train on it.

The problem isn't the test, its that is a public test.

Simon has previously said he has a list of secret prompts (at least one of which he "burned" as a demonstration a while ago). That's what makes it a good test - his commentary on the public test is something of a proxy for non-public tests. This makes it a good benchmark.

andrewstuart 4 hours ago||
The pelican prompt is ridiculous.

Test the LLLM against things you want it to do.

Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking.

Remember these Microsoft interview questions designed to identify the best developers?

"If you could eliminate one U.S. state, which one would it be?"

"How would you move Mount Fuji?"

Absurd interview questions have an air of legitimacy due to the quasi sophisticated justifications put forward for why they are good tests.

Absurd interview questions are not good tests of people or LLMs.

Relevant questions are good tests.

user- 3 hours ago||
The whole point of "AI" is arbitrary task completion. Why isn't a SVG drawing relevant for that?
zahlman 3 hours ago|||
> Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking.

Yes; what's wrong with that?

Do you suppose that it doesn't test those qualities?

simonw 4 hours ago|||
> The pelican prompt is ridiculous

Yes, deliberately so.

It was never intended as a meaningful benchmark. The surprising thing was that for the first ~12 months performance on the stupid pelican benchmark did seem to correspond to the performance of the models on other tasks.

That pattern no longer holds - Fable 5 and GPT-5.6 have both been out-pelicaned by lesser models now.

ErrantX 4 hours ago|||
As I understand it; the point is to ask for an SVG which would demonstrate a conceptual understanding of what is being asked for and that is an important test IMO.

What sufficiently hard, but useful, problem would you ask the model for?

BigTTYGothGF 3 hours ago||
> Test the LLLM against things you want it to do

I agree, it is ridiculous to ask an LLM to replace an artist.

TZubiri 4 hours ago||
https://en.wikipedia.org/wiki/Goodhart%27s_law

"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."

Or the more pop layman version

"When a measure becomes a metric/KPI, it ceases to be a good measure."

Story time, I live in Argentina, and we don't have Big Macs, the main Mc Donald's brand, here, because during the CFK presidency, one of her tactics was to Goodhart economic metrics. Even the informal obscure ones like the [Big Mac Index](https://en.wikipedia.org/wiki/Big_Mac_Index), I don't know the precise details, but the Big Mac ended up being a very cheap item, like 2 or 3 times cheaper than actual menu items, but it was never on the advertised menu, and it also ended up being very small compared to the other burgers, so it wasn't even like a hack, a shrinkflation type of deal.

But hey, anyone who read the Big Mac Index table would never find Argentina at the bottom of that list along with a couple of other countries with bad brands, so the ploy worked. And now we live with the aftershock, the brand never really turned around, other brands with ridiculous names took over it like the McTasty, which makes me sound like that skit from Tarantino's Pulp Fiction.

zahlman 3 hours ago|
How did the president manage to influence McDonalds' local business decisions? And how did that lead to McDonald's pulling out of the country?
Ilya85 3 hours ago||
[flagged]
Ilya85 3 hours ago||
[flagged]
sbseitz 5 hours ago|
I wish I could downvote this for Pelicanmaxxing lmao.
influx 4 hours ago||
Would you prefer the term Pelicangate?
sbseitz 1 hour ago||
Yass!
theandrewbailey 3 hours ago||
We're going to keep maxxmaxxing forever.
sbseitz 1 hour ago||
I believe you are correct!