llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle"
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...4.2266 cents, 38 seconds.
For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.
UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.
And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Aced it, got the job as a senior software engineer.
The interviewers afterwards said "it is SO refreshing to find a software developer who actually knows how to code - never seen such a high performance focused, well built pelican on a bike - you have the skills we need".
That is the best joke I have heard this year. Ready for a stand-up comedy special. Or a song. Superb!
It's like when you ask your average person off the street to draw a house - it'll almost always be square with a triangle roof, one door, and two windows.
In the pelican/bike example, it's probably a bit of a self-perpetuating snowball too. If the earliest examples were bike left-to-right, flat ground, etc. then they are also being scraped up in future LLMs.
There is this scene in the HBO series Westworld where a "host" says some words in sequence which is shown on a display as she says it. Of course, even me thinking of this scene and connecting it to your comment was not original, someone else clearly had the same programming as me.
A medium blog post says
> Pair what with me?” — the moment Maeve (a humanoid android) uttered those words in Westworld (Season 1, Episode 6: “The Adversary”), something clicked. Not for the average viewer, but for me, a STEM educator and AI enthusiast who, just weeks earlier, had read Stephen Wolfram’s seminal essay, What Is ChatGPT Doing … and Why Does It Work?
It's not even that old - but back when it was aired, an AI that can not just string together coherent sentences, but produce coherent reactions in novel, fully unintended contexts, like Maeve was doing there? It was totally a sci-fi premise.
Now we have AIs capable of that and more, and no one bats an eye.
Required sci-fi suspension-of-disbelief in 2017, and then at some point in the last few years we just blew by that one.
Later seasons of the show were much less dramatically satisfying, but also played out the consequences of the science of artificial intelligence demonstrating as a side-effect that human intelligence and free will might have as much of an uncertain foundation as that of machines.
How much data from the Panopticon, how many parameters would it take to train a model that could predict your responses?
https://blog.nawaz.org/posts/2025/Oct/pelican-on-a-bike-rayt...
I plan to update it with more pelicans from all the models released since.
(Spoiler alert: They haven't improved much since then).
> GPT-5.1 Codex
> monstrosity
What are you talking about? That's clearly a sci-fi pelican on a hoverboard (successor of the humble bicycle) wearing a visor. Truly visionary.
The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.
(Why the drivetrain is on the right, I don't know. But most bike parts follow open standards so it's quite entrenched.)
Bicycle frames are not fully symmetric left-right because you need things like a mount point for the derailleur hanger, and optionally affordances to keep the chain off the stays when the wheel is removed.
Those things have to be on the same side as the chain. Bikes designed for disc brakes additionally need a mount point for the brake caliper on the opposite side from the chain.
Additionally, rear wheels are not symmetric: the spokes on the chain side connect to the hub closer to the plane of the rim. That is, they are more perpendicular to the wheel’s rotational axis than spokes on the opposite side (which is why you should always mount a single pannier on the chain side). This asymmetry is to provide space for the gears.
So once the industry decided to put the chain on the ride, you can’t very well make a group set designed for a left chain if you want it to work on the vast majority of frames.
While I'm sure this factors into things for advertisements for bike components, there is also just a general preference that westerners have for left-to-right motion. Not just in bike ads, but all ads with (or suggesting) movement. And also not just ads, but movies where directors believe left-to-right motion is associated with progression and right-to-left motion is regressive.
I bet if it instead had something to do with black widow spiders we'd find that we're most often looking at the bottom of the spider's abdomen, regardless of whatever non-spider-like activity is supplied.
They’re all so close in proportions.
Because they're computers. They don't have an imagination and the ability to create things from whole cloth the way humans do.
Much like a mother pelican, they regurgitate what they've been fed.
Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.
Bird knees bend same way human ones do
I am guessing its not super common, but it happens just so you know.
It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.
Also 3X token use vs. 1.2
Thank you for doing this, I love your benchmark the most!
"Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing.
Or at least they’re not doing it in a plainly obvious manner."Definitely an upgrade over 1.2
I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.
I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.
Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.
its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.
There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering.
I would imagine the labs have some decent ways to produce novel requirements and then actually validate they are met, without the noisiness of implicit human feedback.
That said, the more I think about it, you are right, there's probably also very good ways to extract signal for all these sessions.
I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.
I have to think that's the future, somehow, and I'm really excited about it.
anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.
depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.
the work does nothing or causes net harm.
Well, that's about the same validity as "In Western astrology..." or "in flat earth theory..."
People already started using contributor API, and your input is irrelevant.
That's the opposite of parasitic.
Compare that to Muse spark 1.3
$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)
It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
We understand theoretically they're taking our data, but yeah, that data is vital to the entire business plan of all these companies and WAY more valuable than people are giving credit for.
I checked up on Mistral recently and saw their Claude-alike coding harness is using GLM now, whatever it takes to keep users on their platform and feeding them data.
Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.
This doesn't mean it's not one of the best models available (clearly it is), but that table didn't compare Fable/mythos (unless I missed it?) and OpenAI will be releasing a much more recently trained model (Astra) any day.
So you shouldn't think "wow, Facebook has caught up"
You should think, "wow, Facebook is less than 6 months behind the frontier" and that they're actually creating good models which is going to be good in many ways (price for customers, for one!)
There are downsides too, but I'll discuss those separately somewhere
It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways.
The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).
Stats:
1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)
Muse Spark 1.3 supports Text, Image, Video, File, Audio inputs. We've only started to see models from China include image and video inputs recently.
(It's probably going to be a bunch of repetitive batch jobs like web search that have no training value)