Top
Best
New

Posted by km144 1 day ago

Claude Opus 5.5(www.anthropic.com)
1769 points | 1098 commentspage 11
dom96 1 day ago|
Just updated KillSwitch-Bench with this new model: https://bench.killswitch-lang.org/

It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.

Foobar8568 1 day ago||
I have just switched to 5.5. First mistake was stale environment variable, didn't realize it was replaced, "oh my memory had stall data" and that's it. Second one, a powershell command had the wrong syntax. Great for my first two prompts.
jdthedisciple 1 day ago||
I dare anyone to convince me the benchmarks are not meaningless.

Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?

How would this alleged difference (most likely bs) actually show up in reality?

GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.

enraged_camel 1 day ago|
Ah, so you didn't read the article.

>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

jdthedisciple 1 day ago||
You didn't read my question, bc that excerpt doesn't answer, nor do they demonstrate

> how would this alleged difference (most likely bs) actually show up in reality?

Furthermore: so they admit it's bs but still placate it like its the next biggest thing ever ... alright

All I'm saying is I refuse to buy into it anymore – yet many on here still do, including ... you?

arendtio 1 day ago||
So funny how both OpenAI and Anthropic post outdated pages at the same time. Opus 5.5 has benchmarks against Sol 5.6, and Sol & Luna 6 have their benchmarks against Opus 5.

Why do those labs keep releasing on the same day?!?

herpdyderp 1 day ago||
Probably one gets there first, then the other rush-releases theirs.
throw03172019 1 day ago||
For this exact reason.
breezybottom 1 day ago||
"Where Opus 5.5’s advantage is very clear is efficiency."

Not efficiency in writing, clearly.

doodlesdev 1 day ago||

   > Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5
Big, if true.
toephu2 1 day ago||
When using max effort, I run into context compaction quite a lot. I haven't seen any increase in context window size at all over the past half year (stuck at 1M).

Are the frontier labs even working on this problem?

mnicky 1 day ago|
Why would you use max? It's usually unnecessary and even prone to overthinking. In my experience, since Opus 5 the medium/high is usually enough (until 4.8 I used xhigh, but never max). Even low is quite usable these days..
toephu2 1 day ago||
actually I meant to say xhigh.

even at xhigh I get context compaction quite a bit.

0xMihir 9 hours ago||
yeah, i didn't notice this on 5
More comments...