Top
Best
New

Posted by D2OQZG8l5BI1S06 8 hours ago

Sonnet 5.5(www.anthropic.com)
559 points | 386 commentspage 6
sergdigon 4 hours ago|
Am I getting out of touch or is it becoming kind of confusing what model should be used when? Sure you have tons of benchmarks pareto cost/perf curves etc but at the end of the day when I have a task to give to a model it is not so clear which model and which effort I should choose ... Also benchmark numbers are often reported with max effort but by default effort is medium and based on the pareto curve on this page, Sonnet 5.5 seems more cost efficient than opus only if effort is low or medium!
hank2000 3 hours ago||
username does NOT check out. so confused.
square_usual 7 hours ago||
Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?
simianwords 7 hours ago||
Important to note that lower model + higher reasoning gives a different (not higher) quality of response than higher model + lower reasoning.

Some tasks are reasoning shaped by nature and you can't just throw a big model at it.

quatotor 7 hours ago||
[dead]
rtuin 7 hours ago||
Any benchmarks other than computer use/agentic coding published yet? Curious to compare more broadly with other models
ramish94 8 hours ago||
In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.

Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)

FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)

CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)

Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family

level87 7 hours ago||
This is crazy, what is the point of all these equivalent models?
eli 7 hours ago|||
Those are just 3 particular technical benchmarks. Presumably Opus is a larger model and has greater world knowledge.
salviati 7 hours ago|||
Price going down on each release
bbor 7 hours ago|||
Yup. Recursive self improvement presented in hard numbers.
bigyabai 8 hours ago||
It's long overdue. Sonnet 5 was terrible API value for agentic coding, there were open models like GLM-5.3 Flash that blew it out of the water at 1/20th of the price.

OpenAI and Anthropic's lead is vanishingly small at this point.

TuxSH 7 hours ago|||
> OpenAI and Anthropic's lead is vanishingly small at this point.

Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.

Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.

SubiculumCode 7 hours ago||||
Yeah, I did kind of feel like the step down from Opus 5.5 was so large as to never make it appealing.
bbor 7 hours ago|||
Your takeaway from "Sonnet 5.5 matches and sometimes exceeds the SoTA worldwide" is "their lead is vanishingly small"...?
_fw 7 hours ago||
I still can’t find a place for Sonnet models, I never have.

I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.

Give me the frontier, or give me the cheapest form of good enough.

calumcl 7 hours ago||
There's even less of a place for it considering the Opus price drop as well, I'll still try it but I see no reason to not just do Opus Low/Med instead.

Interested to see if new Haiku gets a big price drop and is comparable to Luna, Haiku is just incredibly out of date with current basement bin pricing.

EMM_386 7 hours ago||
If you're on a Claude plan and have a lot of tasks at the moment that don't require the frontier, Sonnet is a good model to do that since you get more usage out of it.

Sonnet 5 was not a good model though - hopefully Sonnet 5.5 makes the leap that Opus 5.5 did.

limsungkee 7 hours ago||
Yesterday, I realized that Opus 5.5 is cheaper than Sonnet 5. Now I know the reason.
jtrn 7 hours ago||
Here's my purely academic initial impression based on only what they have released from the blog and the system card:

If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.

BUT

It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.

Some of the more interesting things I found from scanning the system card:

- It is the only model tested that shows no preference for rude or polite style.

- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.

- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).

- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.

- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.

- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.

- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).

Clinical behaviour:

Suicide and self-harm handling is reported as weaker in the API because it

It sometimes called a wish to die understandable.

It sometimes validated self-harm as functional.

It sometimes suggested harmful substitute behaviours.

As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."

Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be

gigatexal 2 hours ago||
Sonnet 5 was so horrible I’m afraid of trying 5.5. I’ll stick with Opus 5.5 for now. Seriously the sonnet and opus 5 series models were their vista moment.
popalchemist 2 hours ago|
Increasingly technically skilled. Increasingly more stupid on a human level. I hate talking to Claude more every day.
More comments...