Top
Best
New

Posted by km144 4 hours ago

Claude Opus 5.5(www.anthropic.com)
783 points | 614 commentspage 5
ayhanfuat 4 hours ago|
Looks like Anthropic is starting to give bank reset as well:

> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.

cogythea 4 hours ago||
Interestingly they've changed their approach to usage resets for this release - with previous releases I've had my usage instantly reset, but now in the Claude app I've got a 'Reset for free' button that expires Oct 22, which seems to effectively be a whole new usage window I can activate whenever's convenient
jdmoreira 4 hours ago||
then they copied that from codex because thats exactly how codex works
NielsHarksen 1 hour ago||
Where in your app do you find this button?
madjam002 2 hours ago||
I noticed a big speedup in Opus 5 on Max x20 since about 10 days ago, and I feel like the model has been performing better.

It would be great to know if this was Opus 5.5 or a lesser incremental improvement, as otherwise it's difficult to judge whether Opus 5.5 is expected to be a big improvement.

It's frustrating that there isn't more transparency here.

jdthedisciple 3 hours ago||
I dare anyone to convince me the benchmarks are not meaningless.

Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?

How would this alleged difference (most likely bs) actually show up in reality?

GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.

enraged_camel 3 hours ago|
Ah, so you didn't read the article.

>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

jdthedisciple 1 hour ago||
You didn't read my question, bc that excerpt doesn't answer, nor do they demonstrate

> how would this alleged difference (most likely bs) actually show up in reality?

Furthermore: so they admit it's bs but still placate it like its the next biggest thing ever ... alright

All I'm saying is I refuse to buy into it anymore – yet many on here still do, including ... you?

ryanscio 4 hours ago||
Input $4/MTok and output $20/MTok is a welcome surprise. Cheaper than Opus 5/4.8, Astra 6, Fable 5.
benjiro29 4 hours ago|
The biggest one is the Cache reads going from $0.50 to $0.20 ... Read/Writes dropping by 25% but Cache reads by 60% has a much bigger impact.
dbbk 4 hours ago||
This makes Fable not really make any sense?
re-thc 4 hours ago||
You bet there will be a new Fable soon.
nozzlegear 4 hours ago||
Pacing the frontier btw
re-thc 1 hour ago||
That's a Fable / Myth(o). The name said so.
petesergeant 4 hours ago||
didn't they say Opus 5 was Fable-level too tho? Let's see, I'm at the point where I don't think benchmarks really tell us very much any more. I'd love it to be as strong as Fable, but I'm skeptical about how that will look in practice.
ramoz 3 hours ago||
It crushes Fable on benchmarks and even in the blogs "real-world" studies. But... they are communicating like it ~sometimes~ provides Fable intelligence?

A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??

bitexploder 3 hours ago|
What if the recent Fable intelligence regression was basically just them serving Opus 5.5 until they got it working well?
alpineman 4 hours ago||
So we skipped 5.1, 5.2, 5.3, and 5.4: we really are plateauing
jatins 4 hours ago||
> We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions.

Thank you.

glub 4 hours ago|
System card: https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...
More comments...