Top
Best
New

Posted by km144 1 day ago

Claude Opus 5.5(www.anthropic.com)
1767 points | 1075 commentspage 8
bilater 8 hours ago|
If you're focused on the price drop rather than the increased capabilities of 5.5 you're ngmi
melonpan7 8 hours ago||
Token spend is noticeably less than Opus 5, I can't really tell if it's better though.
34679 1 day ago||
I don't care how good their models get, I won't sign up for one of their plans until they define "X" in their pricing. 5X of this plan, 20X of that plan means nothing when they never tell you what "X" is.

Maybe this model can finally figure it out for them.

johnmlussier 1 day ago||
They killed cyber capabilities so I have to move over to Daybreak on Codex.
ramoz 1 day ago||
It crushes Fable on benchmarks and even in the blogs "real-world" studies. But... they are communicating like it ~sometimes~ provides Fable intelligence?

A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??

bitexploder 1 day ago|
What if the recent Fable intelligence regression was basically just them serving Opus 5.5 until they got it working well?
jatins 1 day ago||
> We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions.

Thank you.

jdlyga 20 hours ago||
Is this less insufferably annoying than Opus 5? That's the question on everyone's minds. Love Opus 4.8 though.
mewse-hn 6 hours ago||
Yeah I've been stuck on 4.8 too, the opus 5 problems seem mostly fixed. It obeys and writes concisely rather than running off, doing its own thing, and emitting a word salad.
wolvoleo 10 hours ago||
Yup same question here
calibas 1 day ago||
> We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in.

We can't test it properly because it knows it's being tested.

johntb86 1 day ago|
Just make it always think it's being tested, and problem solved.
actionfromafar 11 hours ago||
Would it believe that?
jidaigeist 1 day ago||
>Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude.

Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.

andriy_koval 1 day ago||
My bet is anthropic has NN people org who work hard to distill open models in addition to trying to find what other useful materials they can download from shady torrents.
b38484848 1 day ago|||
it's not safe unless it has commitees with orgies with that weird harry potter dude attached
the_gipsy 1 day ago||
"bad actors" boogeyman, and we should trust some tech weasel to do the right thing? Yea we've seen who they really are, once they get a sliver of power.
leothetechguy 15 hours ago|
The Design Language on this announcement page is a perversion of human nature.
More comments...