Top
Best
New

Posted by sfkgtbor 3 hours ago

Claude Haiku 5.5(www.anthropic.com)
488 points | 229 comments
simonw 1 hour ago|
Pelicans riding bicycles for Haiku at the different thinking levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.

The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.

The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:

  llm install llm-anthropic -U                                
  llm anthropic refresh
  llm -m claude-haiku-5.5 'prompt goes here'
EDIT: Here's the Haiku 4.5 pelican from a year ago for comparison, it was terrible: https://simonwillison.net/2025/Oct/15/claude-haiku-45/
ozgung 1 hour ago||
Of course, the sun again. Everyone knows that a pelican can't ride a bicycle without a sun in the frame and can only go right.
accrual 42 minutes ago||
I wonder if we'll start to see pelicans like a mascot of sorts. You could have a pelican pin on your backpack.

> "What's up with the pelican?"

Well you see in the early days of LLMs we wanted a fun way to test new models, and there was this blog, ...

snowram 18 minutes ago||
Will Smith eating spaghetti is the OG benchmark
Jcampuzano2 14 minutes ago|||
I always find the time/token differences between the xhigh and the max effort levels for Claude models absolutely insane.

Even more so, because in a lot of their benchmarks they use the max models. I honestly think I'd rather these labs use their xhigh models as the default for benchmarking instead since I don't think the average person is even using max.

ijidak 55 minutes ago|||
I find it helpful when you post your link that compares the model to other models in the same class or family, or shows progression over time.

The pelicans all start to look the same after a while.

But seeing the comparison to other models by class, family, or historical progression gives an excellent frame of reference.

simonw 16 minutes ago||
Good call, I've edited my comment.

Here's the Haiku 4.5 pelican from a year ago - it sucked in comparison to Haiku 5.5: https://simonwillison.net/2025/Oct/15/claude-haiku-45/

vinni2 1 hour ago||
I thought Anthropic models didn’t generate images.
simonw 1 hour ago|||
This is SVG, but recent Anthropic models have got extremely good at other forms of visual data.

Here's a Blender model I had Claude Opus 5.5 create: https://tools.simonwillison.net/blender-viewer?url=https%3A%...

And here's some animated pixel art by Opus 5.5: https://tools.simonwillison.net/kakapo-party

And some Monkey Island style music (Opus can compose music too): https://tools.simonwillison.net/scrimshaw-jukebox

Anthropic's models do all of this by outputting code. GPT-6 Astra has similar capabilities - I got this Blender model using that: https://tools.simonwillison.net/blender-viewer?url=https%3A%...

lastdong 27 minutes ago|||
This is great! Love the pixel art and tunes.
hazelnut 57 minutes ago||||
Tried it with GPT-6 Astra with Ultra but the outcome was underwhelming with Blender. Maybe it was my prompting ¯\_(ツ)_/¯
jansan 1 hour ago|||
They are really good at generating artifacts, which are windows within the replies containing all kind of visualization, often interactive.

They are still not great at SVG. I just asked Opus and Fable to add a background to an SVG and the results were, well, not great.

mayli 1 hour ago||
SVG is hard.
vunderba 1 hour ago||||
I’ve been playing around with Opus 5.5 which has made a big leap over previous generations in its ability to use a simple drawing-instruction prompt to generate images.

This creates Sierra AGI-style adventure game scenes painted live from simple Turtle-esque drawing instructions so you can basically provide it an empty canvas and then position text labels on the canvas where you want certain things (tavern, oak tree, etc) and it will generate a custom script for rendering them in a EGA graphics style.

https://kq-styles.specr.net

abirch 1 hour ago||||
They generate svg. You can paste in pngs and they'll convert them to svg with varying degrees of success.
sixtyj 1 hour ago||||
They don’t do raster images.
1ucky 1 hour ago||||
Those are SVGs not images.
jonshariat 1 hour ago||||
SVG is code
fakedang 1 hour ago||||
They're SVGs
sixothree 1 hour ago||||
I've created multiple videos using Claude Code, including music and speech. It generates python which in turn generates frame PNGs that it runs through ffmpeg.

Please don't judge me too harshly for this particular poop video. But here is an example of something 100% generated with claude prompts only.

https://www.youtube.com/watch?v=2EqMplbt0gU

zyberzero 38 minutes ago|||
To clarify the ”100%” part - the Python script generated the video output, and you did nothing? No video edit at all? Then I think it is impressive! Are you able to share the prompts you used?
sixothree 14 minutes ago||
Source code is linked from the video! Scan the QR code. I will try to /resume tonight and give you some prompts.
Bluestein 17 minutes ago|||
The Purple Screen of Death at the end :)
ghoshbishakh 1 hour ago|||
Bruh. Svg. It is like drawing something with geometric shapes which are represented using equations.
minimaxir 3 hours ago||
Pricing is...a bit weird.

    Input
    $0.10 / MTok for prompts up to 100,000 tokens
    $0.50 / MTok for prompts over 100,000 tokens

    Output 
    $0.50 / MTok for prompts up to 100,000 tokens
    $2.50 / MTok for prompts over 100,000 tokens
100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.

In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])

dannyw 3 hours ago||
Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here.

For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.

These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.

In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.

flockonus 20 minutes ago|||
> it’s really impressive how much intelligence per dollar has grown in just a few short months.

Open weights models giving a distant salute from afar

WinstonSmith84 2 hours ago||||
noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts. It's hardly competitive ...
JacobAsmuth 31 minutes ago|||
The benchmarks are very long form logic, knowledge, and coding tasks though. I'm very interested in Haiku 5.5's performance on ObviousBench where Luna 6 is currently SotA.
Tiberium 3 hours ago|||
There's also a tokenizer efficiency difference: modern Claude's 100K tokens are about ~60-65K modern GPT tokens, so in reality the Luna cutoff is much further away than the Haiku one.

You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc

AtNightWeCode 2 hours ago||
> ...this tokenizer, the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5.

So, it is might be even worse.

Tiberium 2 hours ago||
No, it's just Haiku 4.5 is so old that it predates the new Claude tokenizer change in Claude 4.7+
Eridrus 3 hours ago|||
It's actually existing flat per-token pricing that is weird.

Neither encode nor decode are linear in compute, so providers need to price for average expected length.

This is just getting closer to the true cost of generating tokens.

hgoel 42 minutes ago|||
Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.
sebzim4500 2 hours ago||||
Flat pricing is weird too but jumping up 5x at one cutoff is surprising in the other direction IMO
foota 1 hour ago|||
My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees.
tr4656 3 hours ago|||
Luna does as well, but just at a higher limit.

From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.

minimaxir 3 hours ago|||
Huh, that disclaimer is on the model page (https://developers.openai.com/api/docs/models/gpt-6-luna) but not the pricing page. Annoying.

Fixed.

tripleee 2 hours ago|||
So even at the 1.5x/2x rate luna is still half the price of this. Weird pricing strategy from Anthropic. I'm sticking with Luna if I don't need a super smart model
usef- 1 hour ago||
You're judging purely by token cost I assume, not cost per completed task?

The benchmark in the article showed it as lower per completed task than luna, but I guess we'll find out how representative that is. Anthropic has generally been fairly honest in their benchmarking though.

tripleee 1 hour ago||
yes, that's true. I should be looking at the $/completed task
HarHarVeryFunny 3 hours ago|||
Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor.

For this application 100K token input is plenty.

Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.

martianvoid 3 hours ago||
I think the 2.5 times cost but actually pays off in terms of intelligence compared to jev and the general capability of using it beyond classification
HarHarVeryFunny 2 hours ago||
The classification performance remains to be seen, but presumably we'll soon start to see classification benchmarks.

For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost.

I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.

alexchamberlain 2 hours ago|||
Isn't it less than a year since Claude models went from 100k token limit to 1M limit? Don't get me wrong - my main agent normally gets to 25% or so before I clear it these days, but as a subagent, doing research or summarisation, I don't think 100k is "absurdly low".
mkotlikov 53 minutes ago|||
If you look at how different reasoning levels can easily exceed task cost of sonnet 5.5 you will see that you will basically never fall into that under 100,000 token threshold. I mean maybe you can choose low and do a basic summary task, but then you could choose something much cheaper instead. I don't know what Anthropic is thinking with its dumber models.
port3000 3 hours ago|||
They are targeting businesses/API use for fast decision making and agent integration. Plus they now need to be competitive with Jev-type models in that space.
cogman10 49 minutes ago||
I think they are also trying to make sure Deepseek and other chinese models don't eat their lunch. They need something price competitive.
mnicky 2 hours ago|||
You could also use it as a subagent prompted eg by Sonnet/Opus orchestrator agent and for many agentic workflows significant part of the dispatched tasks might be under 100k budget.
giancarlostoro 3 hours ago|||
I with they'd give Haiku like 400k tokens roughly, I think between 400k or even 600k tokens is a sweet spot, but Haiku is basically designed to be for small edits is my understanding, but it sucks because any time I ask Opus to "try" letting Haiku do the work, it just falls apart and Opus comes back and tells me it switched to Sonnet (even before Sonnet finally jumped up to 5.x).

I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.

insanitybit 3 hours ago|||
I mostly use Haiku for really, really basic stuff, never for actual engaging work. I've used it for first-pass analysis to triage bugs, for example - all it does is related N bugs together to see if any potentially relate. Then I have Sonnet investigate further.
sixtyj 1 hour ago|||
Chatbot could be < 100k tokens.
enraged_camel 3 hours ago|||
>> 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents

Your vibes don't appear to be supported by facts. From the announcement:

>> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.

Philpax 3 hours ago|||
People weren't using Haiku 4.5 for agents before. 5.5 is good enough that it might be.
StilesCrisis 2 hours ago|||
Haiku 4.5 users were using it for Kleenex requests because that was the best it could do.
enraged_camel 1 hour ago||
Not really. We use Haiku 4.5 to turn users' natural language queries and requests into fairly complex structured specs for interior design and construction. It has near perfect accuracy.
dotancohen 25 minutes ago||
How many examples are in your prompt? How large is that prompt? Or do you have some other way of tuning the output?

I'm asking to learn for a similar project, not to discount anything you're saying.

AustinDev 3 hours ago|||
encode and decode tok/s which is ($/s) when it comes to pricing drops heavily above 100k tokens.

There are plenty of workflows like translations where you'd easily be under the cap.

system2 3 hours ago|||
Who in their right mind would use haiku while Mimo or GLM cost 10% of what they are charging with much smarter models?
wyrdcurt 2 hours ago|||
Some people/organizations are ideologically opposed to using Chinese models. Not me, I use GLM-5.3-Flash for almost everything (the subscription-subsidized pricing on a legacy Z.ai plan makes it the best value model by a wide margin), along with some MiMo and DeepSeek. Still, I use Luna for certain tasks where speed is more valuable than performance; I can see this new Haiku displacing Luna for those. If you mean Haiku 4.5 though I agree, that model was a waste of time and money.
pimeys 4 minutes ago||
Luna is not really the fastest. You need to use it in high/max to get the good output for what it is good for: summarizing. And that is already close to two minutes per task...
mrngld 2 hours ago||||
That's not what any benchmarks that look at cost per task or similar says in terms of cost. The Chinese models, generally speaking, might be cheaper per token but need a lot more tokens to get there.
usef- 1 hour ago||||
On subscription pricing a $20 Anthropic subscription gives >$500 equivalent tokens, which is not so different, and you get smarter models. API pricing has decent margins.

And Opus 5.5 is really good.

user43928 3 hours ago||||
Presumably everyone who doesn't bother integrating a third party API key into their harness, which would probably be most of the Claude Code users.
pkulak 1 hour ago||||
Where do you get this 10% number? Checking providers I know/respect, and GLM 5.3 flash is $0.15/m. Haiku is $0.10/m.
ray_kay777 1 hour ago||||
People who are stuck using Bedrock in-geo due to their company policy (me).
skeledrew 2 hours ago||||
Well, unless you're using OpenCode Go, it's per-token costs (even if already super low), while Haiku falls under the Claude sub. It's just more straight forward and you aren't feeling a "loss" with the sub.
aesthesia 2 hours ago|||
There really aren't any models at 10% of the price of Luna or Haiku.
esafak 3 hours ago|||
It's their creative way of 'matching' Luna's prices.
j45 3 hours ago||
It could be to incentivize people to not be lazy users of tokens.
charlesabarnes 3 hours ago||
> Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users

This is a very big benefit for me. I can now ship actual ai enhanced features behind my subscription without paying extra or fully relying on on-device models. I do worry that this is to soften the blow for user-unfriendly changes

thepasch 2 hours ago||
This is them sneaking in taking the Claude Agent SDK (claude -p) off of subscription plans through the back door along with a model release. They previously wanted to do this in June, but backpedaled after huge backlash:

https://support.claude.com/en/articles/15036540-use-the-clau...

tekacs 1 hour ago|||
I hope people notice again that this is happening this time around.

Being forced through the non-OSS Claude Code with all of its quirks and issues is... such an exhausting use of force by Anthropic.

To the extent that you _can_ choose to disable telemetry and training on your traces in CC, it's not all that obvious what they gain by crippling your ability to use the subscription with other – better – tools.

It's also remarkable that it's coincident with OpenAI adding "Sign in with OpenAI", so that you can use your tokens with other tools.

eli 1 hour ago||||
The page does not say anything about changing the way Agent SDK bills. I just tested Agent SDK and nothing has changed (yet).

You might be right and they will change this in the future, but that's speculative

thepasch 1 hour ago||
> Claude Max and Team plans now include monthly API credits, which cover the Claude Agent SDK, the Claude API, and Claude Managed Agents.

This text has replaced the entirety of the page called "Use the Claude Agent SDK with your Claude plan."

eli 1 hour ago||
Yes. Previously the page was all about how they were going to start charging for Agent SDK use with a banner on the top saying that, actually, they weren't going to do that.
thepasch 1 hour ago||
...yes, a banner which has also now disappeared and been replaced, with the explicit mention that API credits "cover the Claude Agent SDK"?

What more do you need?

eli 1 hour ago||
Well, it doesn't currently work that way on the latest SDK. If there's a change coming, it hasn't happened yet.
thepasch 37 minutes ago||
The monthly credit allocation hasn't rolled out yet either, so as of right now, we're effectively at the status quo. I'd expect the billing change to land once you can actually collect your Claude Console account.
vmg12 1 hour ago||||
These are api tokens, you can build a business with them using any harness.
tekacs 1 hour ago||
Yes, but they're wildly lower in value than the corresponding subscription usage.
sanex 2 hours ago||||
Those mfers. I'm using this for work! I use my work teams plan with pi so I can do all kinds of custom workflows that I can't in Claude Code. Time to convince management I need OpenAI instead.
stsch 2 hours ago||
Time to convince management (and yourself) to build some skills. :)
sanex 1 hour ago||
Spent many years building skills, I'm just working on a different level now.
sambaumann 2 hours ago||||
Even after the June changes there was some allowance to use agent SDK on the pro plan. This will move me to codex tomorrow if agent SDK is really blocked on pro
martinald 2 hours ago|||
Do we know if claude -p is now drawing from this API usage?
ziga 1 hour ago|||
Not yet, as of version 2.1.293. But I suspect this is coming next.
thepasch 2 hours ago|||
That's what the Agent SDK is, according to this help page:

https://code.claude.com/docs/en/headless

So, as written, yes.

geek_at 3 hours ago|||
This is literally for you to get tangled in their api and when they stop giving you the allowance they hope you will just continue to pay
losvedir 2 hours ago|||
Nah, it's pretty trivial to switch providers (especially with Claude's help, ha).

This is more to encourage people to try out adding AI into their product, which is a totally different flow and experience from using AI to build the product.

eli 1 hour ago||
Or to discourage people from using cheap subscription tokens as part of automated workflows
enraged_camel 2 hours ago||||
What does "tangled in their api" mean? Switching is pretty easy.
charcircuit 2 hours ago||
Not really, you have to fiddle with generating api keys and setting environment variables. Meanwhile with Anthropic it will just start charging you API prices for the tokens you are generating without even a single warning.
enraged_camel 2 hours ago||
>> Not really, you have to fiddle with generating api keys and setting environment variables.

That's 5-15 minutes of work at most. Not exactly the type of lock-in the parent is implying.

charcircuit 2 hours ago||
The user could have always done that regardless of if the user has the option to be charged API rates on or off.
tomjen3 2 hours ago|||
That's an old tactic for an old world. You only need, what, half an hour with your agent of choice to write you out of that?
Topfi 2 hours ago|||
This is massive. So on top of the regular usage, we now have USD 200,- to freely use via the API however we please, even resell? That is a statement, even knowing that inference does not cost them nearly as much as they charge, this is very developer-friendly. Does some minor de-risking for testing concepts. Terms seem to be reasonable [0].

Of course, they don't do this out of pure kindness, but I really struggle to see a negative for subscribers already using a Claude Max subscription, especially given changing to another model is essentially frictionless via OpenRouter.

Compared with "Sign in via OpenAI" which they just announced, this is far less lock-in for anyone hosting services but less interesting for users of said services. With Anthropics approach, you can just use the allowance on your users however you see fit along with any other models and once it's used up, you can still just decide not to use their models for the remainder. With users bringing their tokens meanwhile, there is less flexibility in terms of switching for you, though might be cheaper for users.

Both interesting, each approaching this from a very different direction, each having their own trade-offs. On the OpenAI front, will be interesting whether developers can set specific temp, reasoning budgets, etc. for such "provided tokens" or whether OpenAI exposes that only via the actual API.

[0] https://www.anthropic.com/legal/credit-terms

tech234a 3 hours ago||
OpenAI will probably add this to their plans within a week
alasano 2 hours ago||
With OpenAI you can just use Oauth and get a token to use your subscription.

Anthropic isn't even close to being this useful.

Iolaum 2 hours ago||
Biggest reason for an OAI subscription instead of Ant imo.

Biggest loss is that Ant models look like they are genuinely better.

matsz 2 hours ago||
> Biggest loss is that Ant models look like they are genuinely better.

This changes on a weekly basis, I ended up with subscriptions to most of the providers (except for X.ai).

simonw 2 hours ago||
My complaint about Haiku 4.5 was that it was 10x the price of GPT-6 Luna.

> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens

Haiku and Luna now have the exact same price up to 100,000 tokens. Luna is now cheaper for anything after 100,000 tokens, even after Luna's own price increases at 270,000 it's still less than Haiku.

So it sounds like they've directly addressed that problem. Their self-reported benchmarks are all higher than Luna too.

mrbungie 1 hour ago|
Yep. I was looking at the prices of lower tier models a few weeks ago for zero/few shot tasks (pre Jev) and Haiku rates just didn't make sense at all. I ended up using 5.6-luna.

Good to know that is going back to being an actual option from perf/price perspective.

jjcm 3 hours ago||
Ran image -> html tests for this. I was curious if this smaller model was good enough for complex UI. It was not.

Haiku 5.5: https://html.non.io/lcars-haiku-5.5/

Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5

Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...

One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.

thefourthchime 2 hours ago||
Pac-Man Bench:

Considering the price, no model comes close to being as good as this. However, it did take an extremely long time.

TIME 19m COST $0.16 https://jonclegg.github.io/pacman-bakeoff/#claude-haiku-5-5

All results: https://jonclegg.github.io/pacman-bakeoff/

myzie 2 hours ago|||
Interesting that you have gpt-6-luna at $0.01 vs. claude-haiku-5-5 at $0.16 for this task. I see the score disparity though and I played them briefly. My takeaway from this is that the choice between Luna and Haiku 5.5 may remain nuanced. Luna may be a lot cheaper still and good enough for some jobs. Is that your read of the results?
thefourthchime 53 minutes ago||
Actually, I misspoke. At least as far as Pac-Man Bench, Luna does about as good of a job. The ghost logic's not quite as good, but it also makes a map that doesn't have nonsensical sections in it. So maybe call it a wash.
myzie 34 minutes ago||
Yeah, I was mainly thinking about how much cheaper Luna appeared to be in this case.
rpcope1 1 hour ago||||
Something is not right there. DSv4.1 flash shows $1.89 for tens of thousands of tokens? What am I missing?
onlyrealcuzzo 2 hours ago|||
How have you avoided being sued by Namco?
thefourthchime 2 hours ago||
I'm pretty sure they'll never see this. It's pretty much impossible for anything you do you build nowadays to get noticed anyways.
sparklingmango 1 hour ago|||
> likely best used for small subagent tasks / tightly scoped work.

Hasn't this always been the case with Haiku?

saretup 2 hours ago|||
To be fair, you're making it compete with the best public LLM right now that's 2 size/price tiers above it.
jjcm 2 hours ago||
Sure, but presumably Haiku was distilled from the same training data. Part of this is seeing how much the capabilities degrade as their model size goes down.
BrokenCogs 2 hours ago||
Neither of these look "good" to me. There is so much visual noise on the page, like someone turned the "AI Slop" dial to 11. In fact I prefer the simpler design Haiku made.
twostorytower 2 hours ago|||
It's not really about whether the design looks good. It's about if the model can take the design given to it and replicate it in code. Opus 5.5 matches the designs almost to the pixel. Haiku built something else entirely.
BrokenCogs 2 hours ago||
I guess I'm giving GP feedback about their product diffui.ai, not really about Opus' performance.
jjcm 2 hours ago||||
Totally fair, but I'd encourage you not to look at the design so much as the task. This was a design that's part of a benchmark test suite specifically for image->html conversion. The dense visual noise / complexity / flowing svg shapes are things that most LLMs have trouble with.

It's meant to be a good test, not a good design.

FranzFerdiNaN 2 hours ago|||
It’s not really AI slop, it’s how most modern SAAS websites look like.
seaal 3 hours ago||
The monthly API credits for Max plan seems fantastic, especially considering Haiku pricing. Being able to actually use my Claude plan for other harnesses and use-cases on top of regular CC usage is everything I wanted.

Anthropic has really been doing all the right things in the past few weeks, while OpenAI continues to fumble the bag.

thepasch 3 hours ago||
Note that this is Anthropic Trojan-Horsing the previously announced June change in with a model release, where the Claude Agent SDK can no longer be used with Claude subscriptions and is now billed with API credits only.

https://support.claude.com/en/articles/15036540-use-the-clau...

InsideOutSanta 2 hours ago||
Ah, that sucks. I'm using Paseo to run Claude Code; I guess that just got a whole lot more complicated.
0gs 3 hours ago|||
yeah totally agree. esp how efficient it can be to have a subscription quota-paid orch spin up a bunch of API agents, this is kind of like free money to encourage what was already an easy way to save money (via batch pricing)
skeledrew 2 hours ago||
> Being able to actually use my Claude plan for other harnesses

Wait what? This has gotten their blessing?

neucoas 1 hour ago|||
You could always use Claude models on other harnesses via API... just not via subscription. Now they give you $100 worth of API tokens to use on opencode or Pi. Which is better, but still not the same as OpenAI were you can use the subscription on Pi without problems.
copperx 55 minutes ago|||
Absolutely not.
bouk 2 hours ago||
This is great! Been using GPT 6 Luna for decompiling my childhood favorite game (Age of Mythology) and this means I can throw Haiku into the mix as well. 17352/21965 functions matched so far...
WASDx 9 minutes ago||
How do you validate the functions are correct? I did something similar, letting it (mostly deepseek 4.1) translate from assembly to C but it commonly made mistakes, some really hard to discover and fix.
gizmodo59 2 hours ago|||
can you share more details? was this very involved or asking codex/claude/open code with a 1 shot like approach?
bouk 2 hours ago||
I'll write a blogpost when I actually have it working, but basically I gave the game .msi installer to claude opus 5.5 and said to read these blogs:

  - https://blog.chrislewis.au/using-coding-agents-to-decompile-nintendo-64-games/
  - https://blog.chrislewis.au/the-long-tail-of-llm-assisted-decompilation/
And to setup a harness that will decompile the game and start doing a matching decompilation of every function. It set up a bunch of tooling and started a service in the background to do this actual decompilation campaign. I put some instructions into the main opus chat now and then to e.g. add automatic git pushing including a nice svg chart of progress and to switch model strategies here and there i.e. to do a first pass with a cheap model and then switch to opus/sol if the small model can't solve it.

I could now one-shot a new game, yeah.

supersour 2 hours ago||
Maybe a Show HN? I would be quite interested in seeing the results of this project
anthonypasq 2 hours ago||
i love Age of Mythology, but why did you feel the need to decompile it? Its got a great world editor if you were trying to "mod" it.
bouk 2 hours ago||
I want to get the original (Age of Mythology Gold Edition) running natively on macOS and then port it to WASM to run it on the web so I can easily play it with friends
MisterMunchkin 1 hour ago||
Intriguing, I wonder how far you could go with turning games into websites.

Like could total war become a browser game?

steveklabnik 38 minutes ago|||
I saw recently that someone had ported Halo CE to the web and had 1024 players in Blood Gulch.
vunderba 1 hour ago||||
Probably. There have been dozens of examples of taking old games (Crazy Taxi, Super Monkey Ball, Quake, etc) and making WASM browser equivalents using AI to decompile them just on "Show HN" alone.

They often ship the original assets in a somewhat brazen disregard for basic copyright law even when the games are still for sale on places like GOG though.

kro 1 hour ago|||
Last time I gave that a try (without LLM assistance though) it was really hard as games DirectX calls cannot simply be glued to WebGL so performance was bad.
wyrdcurt 3 hours ago||
About time Anthropic released a competitive cheap model. Haiku 4.5 has been too expensive compared to its performance for months now (in fact I don't remember being too impressed even when it was released). This one actually looks worth using in some scenarios. If it's really as much of a step up from Luna as the benchmarks they've shown indicate, it'll probably replace Luna in my workflows. 100k tokens is a pretty low threshold before the price goes up, but I tend to use these smaller models for smaller tasks anyway.
matltc 43 minutes ago||
My weekly limit __on a Pro sub__ has not gone over 50% since before the pre-Fable promos, but usage has been pretty much the same from my point of view. Maybe I am holding it right? Anyone else getting this?

As such, I do not need to even reach for Haiku, and 4.5 was so inaccurate that it often cost more to do so in the past. Sonnet 5.5/low has been good for this kind of thing, and i didn't even touch thinking tokens or any of that. Opus 5.5 low for questions/repros, medium for implementation, basically never reaching for anything above that anymore. 5.5 has been great, so I'll try Haiku, but don't see myself going out of my way to integrate it.

pkulak 36 minutes ago||
I'm excited for API use. I run some agents, mostly on Luna 6 right now. It's just tool use, web browsing, etc, so something dirt cheap, but also not super dumb, is much appreciated. Having a Luna competitor is nice.
coubri 41 minutes ago||
there is no way in hell that im gonna use Haiku too, tho weekly limits become a problem for me in a last couple of month tbh
d1l 2 hours ago|
At work we use haiku 4.5 for a handful of latency sensitive tasks that are fairly simple. It performs well. Just started testing 5.5 as I’ve been anticipating a nice improvement since it was teased. Results so far are trash. Prompt leakage even. And it’s slower. I guess it’s cheap but I think they got the balance wrong on this.
saretup 2 hours ago|
Curious as to why. Haiku 4.5 has been far away from pareto frontier for a long time. Maybe you need to update your prompt for the newer model in your workflow.
More comments...