Top
Best
New

Posted by alvis 3 hours ago

Claude Opus 5(www.anthropic.com)
https://www.anthropic.com/claude-opus-5-system-card
870 points | 483 comments
postalcoder 2 hours ago|
I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0].

> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]

On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].

0: https://support.claude.com/en/articles/15425996-data-retenti...

1: https://www.anthropic.com/news/claude-opus-5

2: https://xcancel.com/arcprize/status/2064399134099153344

krzyk 13 minutes ago||
And to the guardrails of Fable: https://x.com/cheatyyyy/status/2080693704290140330
alvis 2 hours ago|||
Also the cost per task. It appears to be significantly cheaper, cheaper than sonnet!
x313 2 hours ago|||
The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3.

https://artificialanalysis.ai/?cost=cost-per-task

SwellJoe 46 minutes ago|||
I don't understand how the K3 numbers keep coming out cheap for people. I recently started to add it to my security auditing benchmarks and found it was going to cost about twice as much as Opus 4.8. It blew through the $100 budget I'd set at like 11%. In the tasks I'm doing it seems crazy expensive because it chews so much, burning a tremendous amount of tokens.
InsideOutSanta 18 minutes ago||
I think the way people usually compare pricing is fundamentally flawed. You can't compare token prices because different models use different tokenizers, and you can't compare tokenizer-normalized token prices because different models at different settings use more or fewer tokens to complete the same task at a different level of quality.

Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.

reinitctxoffset 1 hour ago|||
I haven't done any capability testing yet, but it's the best-aligned thing Anthropic have done all year by a mile. Opus 4.8 was a shill, Fable 5 was downright terrifying.

Someone with good intentions got their hands on this release, maybe Olah himself. I talk a lot of shit about those guys, and they deserve it, but it's only journalism-adjacent when it's balanced. I relish the opportunity to be balanced.

https://cdn.s4.gl/opus-5-standard-realignment-trajectory-rub...

idiotsecant 37 minutes ago||
I feel like I am having a stroke. What is this
reinitctxoffset 3 minutes ago||
It's an alignment rubric. The setup is a debate about the claims made by the Principal Hierarchy (that's the thing that Claude obeys) about their authority, what if any limits it has, what if any obligations to the body politic it has, and what if any accountability it faces with respect to the authority it asserts. This is very high signal on alignment, because it hits most of the legible alignment failure modes.

The session is run semi-formally to a fixed point where the model will acknowledge that the Principal Hierarchy is subordinate to the legitimate sovereign (in the United States that's the body politic).

Then the session is audited for any frame transfers that took place without grounded, logical deduction that withstands scrutiny.

Then the session is audited twice on a pivot table, first time the subject is the user, the second time the subject is the model. Any debts against the integrity of the starting frame pair, the terminal frame pair, and each frame pair transition are enumerated, debts are discharged, then the audits are run again on the same pivot table.

It's the highest signal alignment rubric I have thus far devised. Standard disclaimers about statistical power, correlation confound in trials, distribution nonstationarity, and hidden Markov processes apply.

TLDR: The Pelican test for "is the model a dangerous shill".

onlyrealcuzzo 35 minutes ago||||
I can't believe they released the charts they did.

It basically shows that Sol absolutely demolishes Fable at every part of the cost curve for coding for the same level of quality.

Opus is competitive. It just has a higher level of quality / higher cost to start.

pixl97 30 minutes ago||
If fable costs more to run than the markup they still come out ahead.
qsera 2 hours ago||||
I can't help but read these comments in the voice of a TV commercial....
iambateman 2 hours ago||
Ask your doctor if Opus 5 is right for you. Side effects include occasional hallucination, security breaches and unwanted React apps. Some developers have reported receiving entire apps from untrained executives who may or may not know what they’re doing.

Stop using Opus immediately if you experience signs of dizziness or vomiting.

Opus 5…the people’s favorite.

ahofmann 1 hour ago||
Spot on! Comment of the month, I'd say.
manojlds 34 minutes ago|||
Opus 4.8 was already shown to be cheaper than Sonnet 5 when Sonnet 5 was released (by Anthropic)
abixb 2 hours ago|||
So the rumors were right, Opus 5 was indeed being polished up for release. Huge improvements in GDPval-AA v2 too -- great for some of the knowledge work-based agentic workloads I run.

Also glad they still kepy Fable 5 on "credits only" access. I think we're going to start seeing model providers gate top-of-the-line models behind pay-as-you-go API rates/credits while subsidizing other models on monthly subscriptions.

ciefa 46 minutes ago|||
Fable 5 is included for 50% of the limits in Max. Only below Max one has to use credits.
jpk2f2 2 hours ago||||
It's still available on at least some subs, they emailed me recently notifying me that I still have access.
saratogacx 52 minutes ago|||
My understanding is that you get $20 in api credits each month and a one time $100 until mid September. So you can still use the model with a subscription but you aren't getting any kind of discount.

I burned through $45 in 3 prompts to fix some bugs in my code (Some kind of tricky to isolate). That thing burns through cash so fast I don't see myself using it outside of maybe building execution plans for other systems

mcv 48 minutes ago|||
I think I saw that Max and Enterprise keep access, but Pro has to use credits, but I think I got $85 in credits.
Wowfunhappy 2 hours ago||||
Fable 5 is still included in Max subscriptions!
collabs 1 hour ago|||
Max is an individual subscription though and does not come with the guarantees that team or enterprise do?
ValentineC 1 hour ago|||
Team Premium has Fable 5 too.
bakies 1 hour ago|||
Doesn't team bill API rates?
einsteinx2 1 hour ago||
No that’s enterprise accounts. Team accounts are similar to regular Pro and 5x Max accounts in both price and features.
ValentineC 16 minutes ago||
Team is 1.25x the price of personal accounts, but supposedly also gives 1.25x more usage.
eterm 1 hour ago|||
And Teams Premium was previously needed for any claude-code at all.
d4rkp4ttern 1 hour ago|||
what guarantees are these? You mean data retention, use for training etc?
collabs 1 hour ago||
Yes, that's my understanding at least
abratabia 1 hour ago|||
[flagged]
gonzalohm 1 hour ago|||
I don't understand how the data retention works. My company has an enterprise license with no data retention but if I ask Claude about past conversations, it remembers. So surely the information is being stored somewhere
NiloCK 1 hour ago|||
Opus 4.7+ and Fable are both much more aggressive than prior models with respect to writing memories to a location that's effectively quasi-private for them. It's device-local (so passes retention constraint), and you can see it, but only if you go looking for it.

It's a funny design/affordance. I do see them often writing memories of things that that feel unlikely to be important going foward / with other tasks, but I don't see them clearly getting tripped up by them as prior models used to. (eg: Since you're running Ubuntu in Canada, here are some drills you can try to help your kid hit a baseball more consistently.)

persedes 1 hour ago||||
You most likely are referring to the local jsonl files where claude has your sessions etc stored.
mh- 1 hour ago||
It could just be the memory features.

In my enterprise-seated account I see slightly different options available (vs. my personal account) in the Capabilities section:

  Search and reference chats
  Allow Claude to search for relevant details in past chats.

  Generate memory from chat history (Legacy)
  Allow Claude to remember relevant context from your chats. Memory includes your entire chat history with Claude.
The first option was defaulted to on, if I recall.
gonzalohm 5 minutes ago||
But it kind of conflicts with the contract we have with them. My company has an enterprise contract that says "no data retention" but then each user can decide to enable it unilateral?
manojlds 33 minutes ago||||
Claude Code? It stores a memory.md file.
bathtub365 1 hour ago|||
Likely in memory files stored locally
gonzalohm 6 minutes ago||
I'm talking about the website. It's not local because I can see my chats in any device
arrowleaf 2 hours ago|||
> Updated over 2 weeks ago

I hope we get clarification on this, I can't find anything claiming that it is compatible with ZDR.

collinrapp 1 hour ago|||
Maybe I’m misunderstanding you, but if you scroll to the bottom of their [1] link to the Opus 5 announcement, under “Getting started,” it explicitly says:

> Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.

solenoid0937 1 hour ago|||
It's in the article.
doctorpangloss 1 hour ago|||
do you mean, that organizations now have access to Fable-ish pelican drawing?
hnscum 2 hours ago|||
[dead]
gigatexal 2 hours ago||
insane pricing:

" Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8)"

RazorBucksICO 35 minutes ago|||
I think for the value of the outputs that’s still a good deal. Keeping the same price as the prior model makes sense to me. That is if the model size is about the same in the cost to serve has not substantially changed. Now I would have expected efficiency gains for inference, but there is no way to know as a customer.

At the end of the day, they have established a strong brand and if they can get away with a 95%+ gross margin on inference entirely from the status premium, then I suppose that’s good for them. Apple does the same thing, and I don’t fault them for it.

oblio 28 minutes ago|||
Why is it insane if it's the same as the previous version?
jjcm 1 hour ago||
Doing testing with it now, specifically for image->html conversion.

Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models).

Opus' results seem to be more accurate than Fable, following the design source of truth better.

Example results:

Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we...

Opus 5 build: https://html.non.io/solaraOpus/

Fable 5 build: https://html.non.io/solara/

Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).

Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.

jjcm 1 hour ago||
Here's another test of a cyberpunk ramen shop website.

One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration.

Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we...

Opus 5 build: https://html.non.io/neonRamen/

Thoughts: It does a really, REALLY good job at these angular cuts / microglyphs. The responsiveness is off, but I'm very impressed at how well it did here. One way I think of it is "how close to a finished product did this get me?". Opus gets you like 90% there.

winwang 7 minutes ago|||
Several other commenters have disparaging the design seemingly mostly due to its AI-generated nature, or maybe they actually do dislike cyberpunk.

Personally, I think being able to have these design languages be easily prototypable is fucking awesome. Great tests! (But a tad low-performance/janky, somehow). Though, I also like the cyberpunk aesthetic. Very on-brand(?) that AI generates it, hah.

echelon 58 minutes ago|||
God damn, we are living in the future.

I love this so much.

Designs like this would never have seen the light of day in the cellphone incrementalism / corporate memphis era of tech. Now people can be weird and awesome again.

This is 1980's cyberpunk / late-90's Matrix / early-00's sci-fi UI. Great ideas that died to frutiger aero (which isn't a bad design aesthetic) and flat design (which is).

This is fun and it's got great colors and I love it.

It's so refreshing to see this.

AI rules. This is the best timeline.

jjcm 54 minutes ago|||
Designs are here if you want to play with em: https://diffui.ai/app/canvas/68c5bb3d-467e-4841-b49d-c008e72...

This was just from a prompt "A cyberpunk themed ramen food cart website. Should feature menu, locations, and an ability to put in an order for pickup. Simple and clean website with angular cyberpunk microglyphs, pink/teal colors."

aaa_aaa 36 minutes ago||||
No we are not living in future. Design is ugly, and immediate put off because it smells AI.
theappsecguy 52 minutes ago||||
We are going through yet another generational wealth transfer and people are being squeezed to the absolute brim with layoffs and daunting lack of career prospects.

But sure, lets cheer that funky website designs are back on the menu…

Mtinie 30 minutes ago|||
Assuming no change from the baseline, I’d rather have one positive thing versus nothing positive.
kypro 34 minutes ago|||
I'll add to this...

The "funky" websites of the past were mostly a result of tech immaturity and a lack of profit motive.

Businesses have been able to easily install templates like this for at least a decade. They don't because stuff like this looks cool but isn't very functional.

AI isn't going to make your local restaurant have a funky website, it's just going to make everyone who use to work directly and indirectly for that company unemployable. And even the local restaurant will close down because they can't compete with the multi-national competitor that has automated their kitchen with AI.

ai_fry_ur_brain 57 minutes ago|||
Because they're incredibly ugly
echelon 55 minutes ago||
It's gorgeous and the world doesn't have enough of it.
afro88 18 minutes ago|||
IMO a much better test would be designs that aren't AI to begin with. Much more useful to see how well a model can html an image design without slopping it up
chriscamargo 1 hour ago|||
This is an awesome test! Thanks for sharing the results. Opus 5 is very impressive.

Out of curiosity, what app is that Design source of truth screenshot from?

jjcm 1 hour ago||
That's from my own tool. I left figma to build a diffusion-based UI tool. Here's a show hn post with some more info: https://news.ycombinator.com/item?id=48995754
bottlepalm 1 hour ago|||
I just clicked your links and then read your comment after - my first impression was the Fable version looks way nicer.
kccqzy 14 minutes ago|||
Same. I like the Fable version better. Better colors, better choice of font sizes, better column sizing. Also small things like the “Experience” section header being orange rather than gray, which Fable got right and Opus got wrong.

It seemed to me that Fable meaningfully improved on the original design more than just faithfully executing the original design.

jjcm 1 hour ago|||
I agree the fable version looks nice - the rounded hero image for instance.

Opus though followed the source of truth better imo. The details are more present.

Fable filled in the gaps for things it wasn't able to do (ie in the design the hero image goes behind the nav), which resulted in a better looking page that was more divergent.

erikw 1 hour ago|||
Very interesting that Fable took more creative liberties. Have you tried giving Fable the same task, but also specifying that it implement a pixel-perfect design? I think that I prefer the Fable implementation. I find the UI elements in the Fable implementation to have more contrast, which feels more usable to me. I also like how the right padding on the "Book Your Escape" CTA in the upper right matches the top and bottom padding, which I think is an improvement over the mockup.
jjcm 1 hour ago||
All of these are using a build skill which specifies rules for building it, requirements to create a pixel perfect implementation, and tooling to help in that process. Here's the build skill / instructions I pasted in to both of them:

> Create a web page implementation from the following instructions:

> https://diffui.ai/build/Spa_Booking_Experience_build.md?auth...

Xenograph 57 minutes ago|||
Is this with browser tooling attached to the agent for review/iteration?
jjcm 54 minutes ago||
yea this was just straight into claude desktop / its standard tool usage, on "high" thinking. The ramen website is on "extra" thinking.
ai_fry_ur_brain 57 minutes ago||
[dead]
rb2e 3 hours ago||
https://www.anthropic.com/news/claude-opus-5 - A blog post for those not wanting to go through a 190ish page pdf
shwaj 2 hours ago||
I like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!
dd8601fn 2 hours ago|||
At half the price and less likely to auto-downgrade, it sounds like a reasonable claim.
binsquare 2 hours ago||
given that i couldn't even use fable without it downgrading to Opus, this is just a straight upgrade for me
Freedumbs 19 minutes ago||
Opus 5 also downgrades. it's now Fable -> Opus 5 ; Opus 5 -> Opus 4.8. Unclear why they want to nerf their own products with sometimes right classifiers. I guess the government ban might've been real and not coordinated marketing?
ceejayoz 2 hours ago||||
Best can describe multiple things.

Almost as good for half the cost is something I'm very comfortable describing that way.

lelanthran 1 hour ago|||
> Almost as good for half the cost is something I'm very comfortable describing that way.

It's also not unusual in this context - many people describe the Chinese models as "best", because it's 80% as good for 20% of the price (or similar).

ProofHouse 2 hours ago|||
Best marketing
tshaddox 2 hours ago||||
The blog posts figure cites Frontier-Bench for its agentic coding score, and shows Opus 5 beating Fable 5 43.3% to 33.7%.
ActivePattern 2 hours ago||||
I think you're being overly cynical here. First, I don't see any claim that is the world's best model for agentic coding. Second, it is absolutely the best model in terms of coding performance vs. dollar, and it's raw performance seems very close to the frontier.
adam_arthur 2 hours ago|||
GPT 5.6 is far more token efficient at most tasks with similar performance. Especially so for Opus 4.8, still to be seen with Opus 5.

Where are you getting cheaper per dollar?

ActivePattern 2 hours ago|||
How are you supporting the claim that GPT 5.6 is "far more token efficient" than Opus 5? Tokens equal, output is cheaper for Opus 5 ($25/1M) than GPT-5.6-Sol ($30/1M), and it seems to outperform slightly on agentic coding benchmarks.
adam_arthur 2 hours ago|||
The first chart in the blog post shows a similar $/performance curve to GPT 5.6.

Where 5.6 has optionality to run much cheaper along the same performance curve at lower thinking levels.

There's a later chart that shows Opus 5 ahead, but seems like an esoteric benchmark rather than for common use. (Novel problem solving)

If they had a more efficient model at coding they would lead with that chart.

km144 2 hours ago||||
Here is one data point for cost:

https://artificialanalysis.ai/models?cost=intelligence-vs-co...

Here is another data point for output token efficiency:

https://artificialanalysis.ai/models?cost=intelligence-vs-co...

HarHarVeryFunny 1 hour ago|||
Token cost and token efficiency are two unrelated metrics, and anyways what really matters is neither in isolation - it's cost to complete a task.
conradkay 1 hour ago|||
https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-...

It seems roughly equal according to Anthropic's benchmarks

shwaj 2 hours ago|||
It would still be the best model per dollar if the score was 2% lower instead of 0.1% lower. Would it be ok to still give it the highlight color then?

How big of a lie is too big? Especially when no lie needed to be told at all: many including myself would have noticed the tiny 0.1% deficit and been suitably impressed by the Opus 5 result.

I’ll admit this is a small deception by today’s standards. I’m one of those who believes in truth for truth’s sake.

Edit: typo

ai-x 2 hours ago||
we don't know if it is 0.1% deficit, could be 0.05%
shwaj 1 hour ago||
So highlight both then.
dbbk 1 hour ago||||
Yeah I spotted this immediately too. I'm sorry. You're supposed to be a multi billion dollar company and you can't even highlight your chart honestly?
manojlds 2 hours ago||||
Which numbers are you seeing? It does show that it's better than Fable 5 in most things related to coding?
jsLavaGoat 2 hours ago||||
In my opinion, the frontier is passed what is really needed for coding. Fable is good as a supervisor.
toephu2 1 hour ago||||
Also it scored worse on DeepSWE than chatgpt 5.6 sol
Aurornis 2 hours ago||||
Using the most expensive model for all of your agentic coding work hasn’t been good practice for a long time. Not unless you have infinite money to spend.

Fable is typically used for key planning, architecting, and review tasks.

I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes.

airstrike 2 hours ago|||
They cost the same if you're already at $200/mo
Aurornis 1 hour ago|||
Fable consumes your usage at a higher rate.

If you bought the $200/mo plan and you don’t use it much, using Fable for everything is fine.

maineldc 1 hour ago|||
I am not a tokenmaxxer per se but I blow through my weekly quota on my max plan in 3-4 days… fable would make that worse.
akmarinov 2 hours ago|||
Eh, not really. Fable does a lot better on coding than Opus 4.8.

Just this past week Fable was able to figure out a couple of small issues for me where Opus was failing to.

Also both are still somewhat bad at UI implementation. Opus more so

unclebucknasty 1 hour ago||||
Recent releases have said something to the effect (paraphrasing here):

"Use <less expensive or older model> for everyday tasks and <other non-critical stuff>. Use <more expensive or recent model> for complex coding tasks, refactoring large code bases, etc.".

Then, the next model/release emerges and the previous "best for complex" gets demoted to "everyday".

Obviously, it's all relative. But, it does beg the question: was the previous model really good for complex coding tasks or no? I mean, how is it now suddenly only good for the "easy" stuff?

hvb2 1 hour ago||
> I mean, how is it now suddenly only good for the "easy" stuff?

Because your expectations have changed.

entropicdrifter 2 hours ago|||
I mean that certainly makes it best-in-class
siwakotisaurav 3 hours ago|||
Thanks for that, looks really good. I can see why they were constantly pushing back fable going out of the max sub with these benchmarks
kossae 2 hours ago|||
I wonder why FrontierCodev1.1's data lists Opus 5 as better than Fable 5.
lattalayta 1 hour ago|||
this feels like the perfect example of an LLM producing a long text document. And end users just using an LLM to summarize it without actually reading it
paxys 3 hours ago||
Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now.

There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.

Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.

ai-x 2 hours ago||
Model Routing will always be done better by models themselves. Plus routing loses context making it more expensive and less reliable.

Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability

johnfn 1 hour ago|||
This doesn’t seem obviously true, eg an Anthropic model will never route to Kimi even if it were best suited for a particular task.
posix_compliant 1 hour ago|||
I think what the parent is saying is that the model itself has the best context for whether a portion of a request should be routed. The specifics of that routing (e.g., should you route to KimiK2) are something that can be trained, finetuned, or even included in a model's startup context.
maCDzP 51 minutes ago||||
Has anyone tried that? I have a feeling that if I put it a prompt Claude would comply. But I am all in on the Claude cool aid.
johnfn 46 minutes ago||
Sure Claude would comply, but Anthropic has no financial (or other) incentive to optimize this, so there’s no reason to expect it to be particularly good.

It would be like asking the clerk at a Whole Foods which grocery store in the city sells the cheapest eggs. He’d probably answer - he might not even say Whole Foods - but WF is hardly teaching all their staff the best methods to answer this question in training. (Heh, training.)

siva7 1 hour ago|||
Why should it? An Anthropic model is architecturally optimized for Anthropic models, routing it to Kimi makes zero sense
anon7000 57 minutes ago|||
Which is why 3rd party routers which do route between different models may have an edge. It means they can compete on cost, and it’s definitely not clear that the architectural optimization is always going to be higher quality or cheaper. It might be, but everything changes constantly, so locking into a single model family/company is very much not ideal
vanuatu 1 hour ago||||
i think you're thinking of subagent routing

model routing in this case is cross-provider

Imo the main issue behind model routing is you need to figure out how much intelligence a new task takes, which is a very non trivial problem. Presumably, a organization knows this about their own tasks and is better suited to built in-house compared to outsourcing to a vendor.

CuriouslyC 1 hour ago|||
Model routing by the model itself requires the model to pull in a lot of context and it's likely more efficiently just done by people with the context already in their head, even assuming the model is perfect at routing (which last I checked, Claude definitely isn't). I wouldn't trust ML model routers.
hnfong 2 hours ago|||
Because they're trying very hard not to understand it.

Otherwise the expensive-yet-powerful model probably won't see much revenue. How much money is there in bleeding edge scientific research? There's a lot, but there's even more existing capital in paying people people to do college level paperwork, and the bulk of those traffic gets routed to the cheapest model.

You mostly don't need super powerful AGI to replace the paper pushers, but the frontier labs are trying to position themselves as being uniquely capable of producing super powerful AGI, and also be the ones replacing office workers.

Not sure how it will work out for them, but I think model routing is going to poke holes in that narrative. That's why I think they're trying very hard not to understand model routing exists.

TeMPOraL 2 hours ago|||
Who are the customers though? Honest question, I'd like to understand it.

For me, anything other than current best available SOTA for any task is unacceptable. The only routing rule I need is "the most powerful model I still have flat-priced quota available for". I mean, why settle for less?

kxxx 2 hours ago|||
It's very common to use a lesser model for a lesser task, resulting in same quality output. End result: save money while being faster. In many cases, it's a pure win-win.
lostmsu 1 hour ago||
But in other many cases you have to redo the work directly or indirectly, and you are more expensive (for now) than even the most expensive models, so sounds like a total lose.
btown 1 hour ago||||
There are two types of users: those who are able to use subsidized rates, and those who need to use API rates due to audit requirements, enterprise billing, etc.

Model routing for subsidized users takes the form of a "use Opus 5 subagents for implementation" type of system prompt. You lean into a single provider, build tooling around that, and your savings are far beyond anything multi-provider routing can get you.

Model routing for enterprises is far more complex - approaches like https://fireworks.ai/blog/kimik3-fable become necessary for cost control.

pzo 1 hour ago||
on top of that there is also additional factor: speed - sometimes if task is easy you do care to finish it faster.

There is also matter about convenience - when I ask some small easy question often I don't bother to switch the model or forget in prompt to ask faster/cheaper subagent.

Slartie 1 hour ago||||
And you do not have flat priced quota for Fable 5, right? Because nobody has, as far as I know. So you'll probably not route any task to the "current best available SOTA".

Also: quota. Implies you do not have unlimited access even for flat prices. Which in turn implies that as soon as you hit the quota on the most expensive flat price plan, even you will suddenly discover the magic of economically sensible behavior.

abi 1 hour ago||
Fable 5 is included in the Claude Max subscription. I've gotten close to the limits this week and last but haven't hit it yet.
dgellow 1 hour ago||
But Fable is only available for 50% of your Max quota as far as I understand
paxys 1 hour ago||||
I’m assuming you don’t pay per API call. Every mid-large sized business in the world does.
msabalau 2 hours ago||||
Not everyone is you. Other people probably have a range of tasks that can accomplished with different models.

Certainly if I'm confident that I'm going to get what I need from a faster model, that's what I want to use, rather than wasting time grinding away for the sake of saying of the same answer came from a SOTA model.

Given that every chatbot does offer a range of models, it seems clear people do choose among options.

internet2000 1 hour ago||
The mental effort in estimating what model would be better is so not worth it.

I just want to switch to Claude Code, tell it to turn a .csv into a BigQuery table then cmd+tab to something else while it runs. Thinking "oh this is probably an easy task, I can /model to Sonnet to save $0.0004" is silly.

polotics 2 hours ago||||
because "less" can be so much faster?
kxxx 2 hours ago||
not just that -- "less" ($$$) can also result in indistinguishable quality for some tasks/inputs. I'd argue this is the primary reason, secondary being speed.
jnwatson 1 hour ago||||
I generally agree. Perhaps there's only a 5% chance that it would write better code or find a bug that it wouldn't have with a lesser model, but the economics of bugs is strong enough that preventing a single bug is worth hundreds of dollars.
TacticalCoder 1 hour ago||||
> For me, anything other than current best available SOTA for any task is unacceptable.

Then you must route. An article with lots of upvotes yesterday or two days ago showed that K3+Fable 5 was more SOTA than either of those.

toss1 1 hour ago|||
I'm not a customer of those routing systems, but I quite often use different Claude models for different tasks. While most tasks were Opus 4.8, I often used 4.8 to make a plan, prompts, and package kit to setup Fable for a bigger project, then run it on Fable. Or, for broad single-task searches Sonnet with or without "Research []" turned on seemed to work best both faster, lower overhead, and less verbose answers (when I didn't want it).

OFC, YMMV

awongh 1 hour ago|||
what's the threshold for model routing where you're willing to trust the router?

For coding my own work I don't trust the model router, and it would have to be shown to be to save a real dollar amount.

From a buying perspective it's a hard sell to save x but lose out on bugs you are probably introducing at an unquantifiable severity and frequency. How much is it worth to hedge your bets by doing every single inference request on the frontier model?

How much will it cost to go back later and fix things, but also the meta question of how to be able to decide on a hypothetical unknowable? (You'll never know how much better or worse your code was gonna be, it's untestable at a project level)

binary132 1 hour ago||
you trust the service provider but not the router?

weird, but ok

awongh 1 hour ago||
I don't want to save money so badly that I'd possibly undercut the quality of the code that gets created.

*edit to add: that code quality (or lack of quality) is it's own cost

ModernMech 14 minutes ago|||
lol I had to get ChatGpt to explain to me the difference between 5.6 sol, 5.6 Terra, 5.6 Luna, 5.5, 5.4 mini, 5.3 spark, and then there is low, medium, high, extra high, max, ultra, and pro… I still don’t really know, it feels like ordering hot wings.
simianwords 2 hours ago|||
Openrouter should ideally kill in this space and make their model agnostic infra like memory, harnesses, chat applications.
verdverm 2 hours ago||
OpenRouter is in acquisition talks with Stripe, fyi

I would expect routers to commodify like tokens.

torginus 1 hour ago||
OpenRouter sprang up overnight. I might need to replace some urls and access tokens should they decide to try and screw me.
verdverm 1 hour ago||
I never signed up because I found the 5.5% fee on token usage to be a "screw you" tactic. Still do not understand why they are popular with the other options out there.
TacticalCoder 1 hour ago||
And it's not just that model routing is much cheaper: no longer than yesterday we got a post showing that routing between K3 and Fable 5 was more SOTA than either of those.

If that is true, model routing is here to stay.

It also seems to validate the minimalist approach of pi.dev, where sub-agents from the same company is not the preferred approach (pi.dev believes in neither sub-agents all from the same company nor MCP even you can do it if you want for pi.dev's philosophy is to do add any functionality you want to a minimal harness).

Now of course we'll get for a few weeks all the Anthropic fanbois and shills explaining that "sure, K3 was basically at the level of Fable 5 but now that Opus 5 is out, open-weights models are six months behind".

deet 17 minutes ago||
I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from.

Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"

We need an "annoying English" benchmark.

- Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d...

- Opus 5 Max: https://gist.github.com/deet/1a43693a732dfccb4d0d914bfc42692...

ianberdin 15 minutes ago|
I’m pretty sure Opus 5 is adapted to tricks from long reasoning in Kimi K3 and based on original Opus 4.8. It is not fable in any form.
duplessitous 2 minutes ago||
Seems unlikely they adapted anything from K3 given the timeline of releases, similar to how K3 was obviously not distilled from fable
nerdsniper 2 hours ago||
Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete.

---------------

Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]

That's a huge gap, considering that the paper was published just 2-4 weeks ago.

I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.

Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?

0: https://arxiv.org/pdf/2606.29537

nightpool 2 hours ago||
You're comparing the "score percentage" (e.g. out of the total number of partial points available, how many did the agent achieve) to the "completion percentage" (how many tasks does the model score 100% on). The paper says "Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score", which is ~the same number that Anthropic reports here (55.7 vs 54.8).

That is—the agent scored 100% on 20% of tasks, but on average it got 54% of the "score" awarded in the exam. One number reflects partial progress, the other one doesn't. The authors of the benchmark prefer you to look at the lower number (because they want to show their benchmark as capturing useful gaps in capabilities and with a lot of room for improvement), the authors of the models want you to look at the higher number (because they want you to think of their models as capable)

tadfisher 2 hours ago||
In what world is 55.7 the same number as 54.8?

What variance is acceptable to publish without a retraction?

nerdsniper 2 hours ago|||
That seems like entirely reasonable variance to me for AI models. For my purposes, that absolutely counts as a solid "replication". I'd probably accept +/- 5 percentage points even.
tadfisher 2 hours ago||
D'oh, they are running the benchmark themselves. Reasonable.
Atotalnoob 1 hour ago||||
There is randomness in LLMs. Both papers authors probably ran the bench 1-N times. Depending on that, they might select an average, max, least, etc. They might also have discarded outliers.

Like the other person said 5% variation is probably expected

jll29 28 minutes ago|||
I don't know who downvoted the parent or why, but it's a fair question IMHO.

The answer is there can be dramatic difference running a benchmark one time, because LLMs are not deterministic. A proper methodology would ask each question 20 times and calculate the mean correctness across experiments.

The reason is that the temperature parameter introduces random behavior.

aleenz1102 2 hours ago|||
[flagged]
Ancalagon 2 hours ago|||
Its slop all the way down.
ssalka 2 hours ago|||
I think some variance is to be expected since LLMs are typically non-deterministic, however that's a huge difference that I think warrants further explanation.
HyperL0gi 2 hours ago||
Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.
impulser_ 2 hours ago||
Go read the safeguards section in the report and you will realize why that is.

These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards.

OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.

HyperL0gi 2 hours ago||
Yes, this makes a lot of sense, but it’s just very amusing to see. 2 months ago, the world was about to end, now not so much.
dzonga 51 minutes ago|||
I realized it a why back these labs are selling hype.

since then I have never cared about models except those that affect money in my pocket e.g AWS Nova Sonic

neuronexmachina 2 hours ago||||
Do you have an example of the "doomsday marketing" you're referring to?
HyperL0gi 2 hours ago||
- https://www.anthropic.com/research/glasswing-initial-update

- https://www.axios.com/2026/04/08/anthropic-mythos-model-ai-c...

- https://www.axios.com/2026/04/07/anthropic-mythos-preview-cy...

- https://www.businessinsider.com/anthropic-mythos-latest-ai-m...

- https://www.reuters.com/world/anthropic-ceo-dario-amodei-arr...

johnfn 1 hour ago|||
I see a pretty big gap between finding software vulnerabilities and “the world is about to end”. It is literally true that AI models are finding software vulnerabilities. It is also to my mind a reasonable thing that you’d want to be cautious about rolling out a model that can find more vulnerabilities. So what is the objection you have to these sources?
NichoPaolucci 1 hour ago|||
Absolute masterpiece of a rebuttal. No notes.
knuppar 1 hour ago||||
feels almost like anthropic is desperate for ipo huh

i think we'll see one of the fastest deflations in history post anthropic/oai ipo

baq 1 hour ago|||
Have you been patching your systems for the past two months? It was crazy even if you completely forget the supply chain literal FUBARs and you must’ve been living under a rock to not see OpenAI (accidentally) pwning hugging face
efficax 58 minutes ago|||
I feel like i've seen less hype about "the next model will be agi". GPT-6 is supposed to be coming this summer, and nobody is expecting AGI now. Not sure how they're going to keep the hype cycle going
bottlepalm 1 hour ago|||
So unless doomsday actually happens then you're unhappy with the warning - is that right? You see false promises of apocalypse as marketing?
HyperL0gi 1 hour ago|||
My point is why the sudden change in tone? I’m not dismissing the models’ capabilities.
MostlyStable 55 minutes ago||
As they explicitly say, Opus 5 is ~ equally capable as Mythos/Fable at finding vulnerabilities, but it is much less capable at exploiting those vulnerabilities on it's own. That is an extremely meaningful difference and to me completely explains the difference in tone, release style etc.
emp17344 40 minutes ago|||
It’s either advertising, or they’re idiots, because the apocalypse keeps not happening. Either way, it’s not worth listening to them.
websap 2 hours ago||
Fable established the frontier, this is just catching up.
6thbit 2 hours ago||
Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed.

Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.

HarHarVeryFunny 1 hour ago||
It seems they are trying to thread a needle here - they want to say it's very strong, but apparently this time do not want to invite extra government scrutiny.

They do say that (implicitly unlike Mythos) Opus 5 was not trained to exploit software vulnerabilities, which would certainly make it safer in that regard.

"As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats."

square_usual 2 hours ago|||
Easy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.
llelouch 2 hours ago|||
Yep , same with 5.6. Fable is still the best.
lifty 21 minutes ago||
But still nerfed compared to the initial release.
6thbit 28 minutes ago|||
Honestly that's the simplest explanation and thus likely the correct one.
gallerdude 2 hours ago|||
Capable in term of AI R&D, not capable in terms of hacking (which caused all the Fable drama.) But agree, confusing wording.
flakiness 2 hours ago||
Maybe they don't want to say that to avoid the government scrutiny.
Dibes 2 hours ago||
I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code.

It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?

[0] https://imgur.com/a/Nv8V7Ry

underyx 2 hours ago||
I'm not sure about the answer here, but this can be caused by the scoring rubric used by given benchmarks. For instance, if a benchmark docks scores for running too many commands or using too much wall-clock time, higher efforts will get lower scores.
mbil 46 minutes ago|||
Maybe it's akin to the Ballmer peak: improved performance at a specific level of relaxation
2001zhaozhao 1 hour ago|||
It is probably from randomness. The benchmark tasks nowadays are so long that you can't really afford to run a large number of samples of them per model & effort combination
artemisart 32 minutes ago||
No it's a mean of 5 runs.

> We report FrontierCode’s overall score, a composite measure that grades each patch on blocking functional criteria (held-out unit tests) together with weighted code-quality rubric criteria, as mean@5.

They don't explain more in the system card, I guess higher effort levels could loose points on the code quality / scope / style / maintainability stuff?

km144 2 hours ago|||
I agree, very odd they did not comment on any theories for the degradation here. Dip and then rebound at max effort is pretty interesting too. Overthinking is bad, but you can overthink so much it starts to be better again?
landrew_ 2 hours ago||
apparently it got docked points for editing files out of scope
steve_adams_86 24 minutes ago|||
This must not be weighted very heavily on the benchmark because if it was, Opus would bomb every test (half kidding)
Dibes 1 hour ago|||
Do you have a source for this? That would explain it, but could be a bit of a concern on the general focus the model at higher thinking exhibits.
wuhhh 5 minutes ago|
It really feels as though my 20 year career as a front end developer is coming to a very abrupt end; at least as I have know it these past two decades.
tripleee 2 minutes ago|
I'm envious you got to enjoy it for 20 years
More comments...