Top
Best
New

Posted by km144 2 hours ago

Claude Opus 5.5(www.anthropic.com)
542 points | 514 commentspage 3
skunkworker 2 hours ago|
At this point I'm convinced they are skipping numbers so soon they will be at or ahead of OpenAI's numbering scheme.

Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.

ekckekcjekfj 2 hours ago|
And how was the Xbox 360 naming choice a “debacle”, exactly?

It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.

I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.

Gander5739 2 hours ago||
Dupe: https://news.ycombinator.com/item?id=49803863 (or vice versa)
tomhow 2 hours ago|
Comments moved thither. Thanks!
km144 2 hours ago|||
Can you fix the link on that post then? I duped because that post links to a diff that tells me nothing about Opus 5.5
tomhow 2 hours ago||
I did that but I recognize that even though your submission was a few minutes later than that one, you posted the better link, and you're also an established account (the other post was from a new/throwaway account), so I've restored this submission and moved the comments back to it to reward you.
meerita 2 hours ago||
As long as it's not as verbose as Opus 5, I am quite happy with a better version that's also less expensive. I will test it tonight. Grok 4.7 was horrible, and for mundane tasks I am relying on DeepSeek Flash 4.1 with great success using OpenCode.
garo-pro 2 hours ago||
Finally confirmation that Haiku was not forgotten and will be coming soon, althouhg I find it quite interesting they skipped 5 and directly skip to 5.5 with all models, including Sonnet which is not super old. I suspect they found something breaking that allows to release this. Recently they struggled with keeping up a 50 % weekly limit increase and now they're putting out 30-40% faster and cheaper models even faster, with much more better benchmarks, a limt reset command and five hour limit increase. It seems more like the opposite and as if they never struggled, thus, I very much believe they found something very effective and new.
aesthesia 30 minutes ago|
Sonnet 5 was released a while before Opus 5, so it's just Haiku that didn't get a 5 release.
tomaskafka 1 hour ago||
Excellent, maybe Anthropic can use it to fix Claude Code Desktop kicking me back to login every week or so, and forgetting whole state (opened windows = the only way of managing active working set) when I sign back in, if it's that good.

Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).

aragornii 2 hours ago||
What I'm mostly interest in is the Communication section. Opus 5 was so convoluted in the way of answering that was really frustrating me.

Instead of instilling confidence, it was overwhelming. Not sure if I'm the only one.

edude03 1 hour ago||
> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1.

Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM

rumblefrog 2 hours ago||
I'm glad they specifically called out the prose issue, I was always pinned to Fable 5.1 because I wanted to avoid the unreadableness of other Anthropic models.
pookieinc 2 hours ago||
“It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”

They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?

randomblock1 2 hours ago||
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
jbellis 2 hours ago||
Anthropic knows that the benchmarks showing Opus 5 better than Fable 5.1 are measuring something that's less than entirely useful.
meric_ 2 hours ago||
Opus does seem like a more powerful coding workhorse based on the benchmarks listed though. Good coding performance, faster and less verbose, cheaper.

Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too

dom96 1 hour ago|
Just updated KillSwitch-Bench with this new model: https://bench.killswitch-lang.org/

It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.

More comments...