Top
Best
New

Posted by km144 2 hours ago

Claude Opus 5.5(www.anthropic.com)
334 points | 400 commentspage 2
yipinwong 55 minutes ago|
I spent about $5 per sentence in my resume using Fable 5.1 (High) to verify accuracy, inconsistency, and edit.

Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.

Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.

copperx 53 minutes ago|
$5 per sentence?
yipinwong 41 minutes ago||
I am sorry, I meant to say I generated STAR out of my resume line, trying to generate STAR, and polish it thus $5.

---

I provided crapton of context for that one resume line. All the work I did, documentations for my justifications, etc.

I initially messed up and came out ot $5, rest of resume used around $4 per line (I used a fresh new session on purpose).

---

As a clarification, $2.2 average for OPUS 5.5 was the same process in a new session, same context, same prompts.

Also adding verification for that Fable 5.1 output in the same sesssion.

somewhatjustin 1 hour ago||
> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

Nice. I was starting to think that Haiku got abandoned.

Sol- 1 hour ago||
Found this announcement interesting since allegedly OpenAI is retiring their Terra tier. I think for everyday work, two models with various thinking efforts seem enough, plus some frontier level model like Fable or Astra to coordinate.
skerit 27 minutes ago|||
Retiring the Terra tier? Their space-inspired lineup has only been out for 2 months, they're already messing with it?
somewhatjustin 1 hour ago||||
I personally use up to 3 models. Fable/Opus for planning, Opus/Sonnet for implementation depending on complexity.

I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.

ricardobeat 1 hour ago|||
Terra lost to Sol and Luna at every cost/performance point, it had no reason to exist.
mchusma 1 hour ago|||
I hope Haiku is Pareto better than Luna/Deepseek, so slashing its price by about 90%.
cesarvarela 1 hour ago||
Claude code still uses it internally.
zuInnp 1 hour ago||
So it Opus performs as well as Fable what is then the selling point of Fable?

All of this starts to feel more like a drug dealer selling their newest stuff.

In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.

And on the way I always have to check my tooling and need to adjust things to get max results.

ieie3366 1 hour ago||
? they will obviously release Fable 5.5 soon(tm). It's same as hardware. The previously top tier product gets obsolete
ACCount39 44 minutes ago||
New generation's "upper-mid tier" offering claims to be almost 1:1 match for the previous gen's "top tier" - in other news, fork found in kitchen.

Now, Anthropic might stall on releasing Fable 5.5, due to the "pacing the frontier" threat-to-humankind management business. If so, Fable 5.1 would remain a niche model for the next bit.

orangecat 1 hour ago|||
All of this starts to feel more like a drug dealer selling their newest stuff.

Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.

glub 1 hour ago|||
Don't forget that Opus 5 was tracking fable on many benchmarks, yet it was borderline unusable for any coding work. My Claude sub usage has been 100% fable, 0% opus 5.

Benchmarks often don't survive contact with reality.

drnick1 1 hour ago|||
That's not my experience at all. Opus is an extremely capable coder on high or xhigh effort. It can read academic papers, implement algorithms from the description in the paper alone and reproduce results without breaking a sweat. This is remarkable because it is pure reasoning on unseen material; in some cases the paper was just published and there wasn't an implementation to learn from in the training data.
cowthulhu 48 minutes ago|||
My experience is that Opus can definitely write decent code, but it incurs tech debt and adds unneeded complexity.
cheikhcheikh 1 hour ago|||
did you actually verify that it's output in those scenarios is good ? in my experience opus has been a disappointment and constantly trailing behind actually solving hard problems versus the OpenAI models. I'll say that both have terrible writing style though.
drnick1 29 minutes ago||
> did you actually verify that it's output in those scenarios is good ?

Yes, in the sense that it reproduced results in the paper or known solutions obtained by other methods. In fact, Opus is very good at checking it's own work in my experience.

arw0n 1 hour ago||||
Opus is fine at coding (for correctness), but horrible at talking about code. I don't really see the defect rate going down when using Astra or Fable 5.1, but they are just more coherent in both how they explaing code/architecture/choices, and how they actually code the thing. With Opus, I'm using smaller models to delete the vast majority of comments and 'clean up' correct code that is too weird.

Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.

boredtofears 1 hour ago|||
Mines pretty much inverted - my colleagues and I noticed almost zero difference between the quality of code in Opus vs Fable. Occasionally I'll switch to Fable for an arduous debugging task but that's about it.
giancarlostoro 1 hour ago|||
Fable should have just been called Opus Primt for Enterprise and sold only to enterprise customers. I don't even use it. I rather just use Opus.
Imustaskforhelp 57 minutes ago||
isn't this what mythos was/is?
notatoad 55 minutes ago|||
typically, when any AI company says a model performs as well as fable, all they're really telling us is that the benchmarks that exist for measuring AI capabilities aren't very good.
quotemstr 1 hour ago||
Big model smell is a real thing. For certain classes of problem, ones you get a feel for but can't easily articulate, a big last-gen model can get you what you're looking for when no quantity of tokens from some ultra-RLed mid-size latest generation model can.
jjcm 1 hour ago||
Image->HTML tests:

Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...

Opus 5.5's output: https://html.non.io/annui-opus/

Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.

For comparison with other drops this week + current #1:

Astra: https://html.non.io/annui/

MiMo: https://html.non.io/annui-mimo/

Grok 4.7: https://html.non.io/Annui-grok/

copperx 51 minutes ago||
By any chance did you tried Deepseek 4 or 4.1 and GLM 5.3 or flash?
jjcm 16 minutes ago||
I've done GLM 5.3 previously here: https://news.ycombinator.com/item?id=49295420

Worth noting though that GLM 5.3 isn't multi-modal, so it doesn't have a vision layer. It is quite clever and hacks around it pretty effectively however. I'm running a deepseek 4 build now and will reply shortly with that.

naet 1 hour ago||
What is your workflow for making these?
jjcm 19 minutes ago||
The designs are outputs from my own site. This has an overview of the process: https://diffui.ai/learn/new-site

The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.

Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.

For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.

throwaway2027 2 hours ago||
After yesterday outage is the new Opus 5.5 load-bearing?
lgessler 1 hour ago||
I should find information about the user's concern instead of just assuming.

The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.

One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.

handfuloflight 2 hours ago|||
It's worth stating why, and depends what seams you're pulling at this sitting.
danw1979 1 hour ago|||
Good point — but I’ll gently push back on that. It’s not an outage, it’s a service degradation.
staticman2 2 hours ago|||
I'm gonna be straight with you—I don't have the evidence to say whether or not it's load bearing.
rich_sasha 2 hours ago|||
Your instinct is basically right, and the research backs it up.
cmrdporcupine 1 hour ago|||
And here's the important part...
danw1979 1 hour ago|||
you win the thread
ThouYS 1 hour ago|||
You're right to bring this up - and this is where it gets interesting
hmokiguess 1 hour ago|||
You're right, this changes everything, and here's why it matters.
cronin101 2 hours ago|||
It certainly _seams_ that way
aoeusnth1 1 hour ago|||
You were right to call that out, and the evidence makes a stronger case than you are stating.
sailfast 1 hour ago|||
Let me verify before I come back to you with an answer that is incorrect.
loopmonster 56 minutes ago|||
That's the sharpest point anyone has made in this thread so far, and it reframes the entire conversation.
carlos-menezes 1 hour ago|||
One thing worth flagging here: 5.5 appears to be a load-bearing seam in the numbering system.
RGS1811 1 hour ago|||
This question is real.
bibimsz 44 minutes ago|||
One pushback: there is no Opus 5.5. You might have meant Opus 5.1, the latest Opus model available.
fghorow 1 hour ago|||
"Danger Will Robinson!"
esafak 1 hour ago||
Wrong century, brother.
tda 28 minutes ago||
[flagged]
abtinf 1 hour ago||
Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).

Edit to address questions below:

ChatGPT supports oauth login.

Exe.dev has it built in. IIRC, pi also has it built in via /login.

mlcruz 44 minutes ago||
What worked well for me was a custom version of Open Web Ui with some customization to spawn an exe.dev instance for each new chat. I can just work on my phone, deploy stuff for development purposes on an easy to share way etc.
polalavik 1 hour ago|||
ya i've been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.
roughly 1 hour ago|||
> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

Can you give more details here? This sounds intriguing.

sidrag22 1 hour ago||
Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can "do it", but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.

So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).

cbg0 1 hour ago|||
How about cheaper? Astra is $10 in $50 out, Opus is $4 in $20 out. Even on a subscription you'll get considerably more usage out of Opus.
qlte 43 minutes ago|||
Per the link someone else posted, the actual difference in $/task is not nearly so stark:

https://artificialanalysis.ai/models/releases/claude-opus-5-...

  Opus 5.5 Medium = $1.34
  GPT-6-Astra High = $1.76
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):

  Opus 5.5 High = $1.82
  GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.

So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.

margorczynski 22 minutes ago||||
In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.
notatoad 18 minutes ago||||
yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn't run out that quickly for me. it's great, but it's on the same tier as fable for me - use it sparingly, only when really necessary.
copperx 48 minutes ago|||
> Even on a subscription you'll get considerably more usage out of Opus.

That's an incredibly bold assumption.

cbg0 39 minutes ago||
It's not, I have a subscription to both and Astra burns usage like crazy.
onlyrealcuzzo 1 hour ago|||
> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

This is news to me. Excited to try it out! Thanks.

nchmy 1 hour ago||
news to me as well. i thought you were forced to use Codex if you wanted their subscription. I completely ignored it because of that. How do we do it?
KeplerBoy 1 hour ago||
With the pi harness it just opens the browser (or gives you a link if you're on headless) and you sign on as usual.
felixgallo 1 hour ago||
If you read the page, Opus is now significantly better than Astra while also being cheaper and having more performance headroom available.
abtinf 1 hour ago|||
I read the page. It seems like a marginal improvement.
ryanscio 1 hour ago|||
Let's wait for independent benchmarks at least
esafak 1 hour ago|||
https://artificialanalysis.ai/models/releases/claude-opus-5-...
felixgallo 1 hour ago|||
the benchmarks provided are already from independent organizations:

Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)

FrontierCode v1.1 - Cognition

CursorBench - Cursor (now SolarBoringSpaceXAI I believe)

GDPVal-AA - Artificial Analysis

AutomationBench - Zapier

Humanity's Last Exam - CAIS and Scale AI

Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute

OSWOrld - XLANG Lab @ the University of Hong Kong

Chartography - Surge AI

sznio 2 hours ago||
I'm more excited by the Haiku 5.5 announcement buried in this post. I'm wondering if we will finally get a decently capable fast model.
booty 1 hour ago||
If you're able to use the OpenAI ecosystem, Luna's price/performance is really good. Almost like "they messed up and accidentally made it too good" good.
copperx 46 minutes ago||
Better than what? Deepseek? GLM? Gemini 3.8?
copperx 47 minutes ago|||
Why are you excited about it? Deepseek is everything Haiku wishes to be and more.
lanyard-textile 1 hour ago|||
Agreed. They've been so quiet about it, and retirement for Haiku 4.5 is right around the forner.
ygouzerh 1 hour ago|||
What are you using Haiku for?
Zambyte 1 hour ago|||
Not the same person but... nothing. Haiku just hasn't been an interesting model for a long time. If you want cheap and fast, there are lots of options that are simultaneously cheaper, faster, and capable than Haiku.
mavamaarten 1 hour ago|||
I use it for executing well-prepared plans sometimes. And for exploring larger codebases.
system2 2 hours ago|||
All I care about is the token price for the API. Haiku cannot get close to GLM or Mimo.
enraged_camel 1 hour ago||
We use Haiku 4.5 inside our product. It continues to be absurdly capable for converting natural language to structured JSON based on a set of fairly complex business rules.
anthonypasq 1 hour ago||
bro why. its literally the most overpriced model in existence right now. i could name about 10 models off the top of my head that would be better and cheaper
enraged_camel 58 minutes ago||
We tried Luna and it scored way lower in our evals. Muse also. We haven't had a chance to test others.
Game_Ender 17 minutes ago||
How much time were you able to put into tuning your prompts? And was it worse on all fronts (cost, latency, accuracy) or just some?
bredren 1 hour ago||
Notes on communication:

"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"

and

"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."

and

"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."

I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.

If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.

mgw 2 hours ago|
They mention "the first model in our new Claude 5.5 family". Obviously that means Fable 5.5, but hopefully also a usable update to Sonnet and Haiku. Sonnet 5 hasn't really had a place in the line up for anyone I feel.

Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.

mudkipdev 1 hour ago||
It does mention sonnet and haiku.
simianwords 1 hour ago||
And not fable lol
enraged_camel 1 hour ago||
At the end of the post they said Sonnet 5.5 and Haiku 5.5 are coming soon.
More comments...