Top
Best
New

Posted by D2OQZG8l5BI1S06 5 hours ago

Sonnet 5.5(www.anthropic.com)
504 points | 338 commentspage 3
onlyrealcuzzo 5 hours ago|
> In our testing, it costs up to 30% less per task than its predecessor.

> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.

This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.

They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.

Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.

I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).

For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.

This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.

Hopefully they release a Haiku that actually has a reason for existing.

fluidcruft 48 minutes ago||
Would Jev-type functionality be a reason to dust off Haiku?
jchw 5 hours ago|||
I always tell coworkers if they're gonna use Claude to just stick to only Opus and Fable. Sonnet is a waste of time that does a bad job at a bad price.

DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.

velcrovan 5 hours ago|||
Sure, but the fact that Opus 5.5 was such a huge leap over Opus 5 (and Fable 5.1 for that matter) means that it's worth revisiting your priors on a new Sonnet.
eli 5 hours ago|||
Hopefully some faster providers will start offering mimo-v2.6-pro because it's cheaper and benchmarks better than Deepseek
jchw 5 hours ago||
Theoretically but I've used DeepSeek V4.1 Flash for several hundred millions of tokens already and it chews through tokens but it is surprisingly good at making it to the end.

MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.

I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.

I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.

Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.

eli 5 hours ago||
I read they identified a training bug and were going to push out an updated release to fix the looping. I really like it overall.

GLM 5.3 Flash is also very good. I think a little smarter and a little more expensive.

jchw 2 hours ago||
I did like GLM 5.3 Flash but it's just way too often I'd run it on some long running task and come back to it repeating the same tokens or tool calls endlessly, just doing nothing. It wasn't unusable, but I couldn't trust it. That's really frustrating and I think new models have to do better not just on benchmark scores but general reliability and user experience as well.

At some point Anthropic and OpenAI models definitely could fall into similar traps so I do think it is a solvable problem and likely not a reflection of the models themselves being bad. In this case it may indeed be a training bug of some kind, but I also suspect mitigations on the inference side are possibly lacking or not effective enough for the open models and their runtimes.

criemen 5 hours ago|||
The latest Haiku release is almost a year old. Clearly they don't care about the small-but-capable part of the market at all.
enraged_camel 5 hours ago||
From TFA:

>> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.

aniceperson 4 hours ago||
This will be interesting. While no one cared about small models in the last few months except for the OSS community, there is a silent small model revolution with gpt luna and jev. Headless/background llm routines are cost-feasible, which will of course lead to exponential usage and cost.

My take on anthropic is that haiku 5.5 has been shelfed for a while since it is predatory against sonnet (see terra 5.6 usage), but openai went kamikaze and they are now forced to release.

Nevertheless, the elephant in the room has grown: will any of the Labs be able to profit if mass adoption lies in the highly crowded small model territory?

https://openrouter.ai/blog/insights/gpt-5-6-discounts-jevons...

criemen 4 hours ago||
> but openai went kamikaze

I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.

aniceperson 3 hours ago||
Of course, my opinion is based on my personal experience + openrouter data that shows stickiness and low terra adoption; with openai confirming by making sol terra, astra sol.

I honestly never saw anyone doing /model gpt mini. I think those models were mostly used for copilot-like products, like those pull request reviews with untasteful dumbness to it (idiotic CodeQL finding -> LLM vomits a "fix" instead of assessing). While Luna seems to be the first model that you can trust to reason in the background, and this is predatory to their own more expensive model.

cbg0 4 hours ago||
> This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.

It's been out for an hour and you've already concluded this?

AM1010101 4 hours ago||
For me I would like to pair this with Opus 5.5 as orchestrater and use Sonnet as a sub agent. Therefore I want it to be fast when on low or medium and not break the bank.

On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.

If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.

mchusma 2 hours ago||
I feel like sonnet is priced too close to opus right now. If Sonnet 5.5 were half its current price it would make sense to use. At its current prices, I won't use it in applications (I would use cheaper models) and I won't use it in my subscriptions ( just use Opus instead). At least that is my initial reaction.
solenoid0937 5 hours ago||
Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.

At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.

boc 3 hours ago||
I was talking about this with a friend this weekend. We both work in the field and test new models within minutes of them being released. We both immediately clocked Opus 5.5 as being cracked within the first hour. Went on HN and the launch announcement was full of people whining and pointing at cost/token charts vs Chinese models. It was like the upside-down world.

We were both sad that HN has become a negative signal news source on AI lately - you're much more likely to be misled by this website in 2026 on the topic of frontier AI. If you're reading this comment, you should do your own research vs trusting the "Astra is 1000% the best" or "Deepseek is the $/tk KING" comments swarming these announcement posts.

rfgplk 5 hours ago|||
Astra is still the uncontested #1 code generator.
solenoid0937 5 hours ago|||
Astra is amazing, I love it.
dude250711 4 hours ago|||
Yeah, especially coupled with Opus for alternative reviews. A massive token burn though.
ricardobeat 5 hours ago||
I mean, they worked really hard for this. Back in February everybody loved them.
solenoid0937 5 hours ago||
I think all the positive people have just stopped commenting.

The difference in perception for Opus 5.5 on HN vs the real world is what convinced me HN is totally detached from reality.

tombert 4 hours ago||
I like that "alignment on safety" appears to mean, at least for anything I've been doing, that they won't violate Microsoft's terms of service. I even had it pushing back on me activating an LTSC key on Windows because LTSC keys are "often purchased on a gray market and violate Microsoft's TOS".
aniceperson 4 hours ago|
I saw that with corporate software too. What works is creating a skill with the task steps, it fades its initial reasoning. (I am not talking about observer safe guards, but the safety RTL).
sajithdilshan 4 hours ago||
I use Claude Code everyday for work and the main model I use is Opus (For planning, breaking down tasks, writing tickets, implementation, etc.) and Haiku for running tests. Honestly have no idea what is the use case for Sonnet
ricericerice 3 hours ago|
my feeling is you're most likely wasting money using Opus for implementation. The plan and task breakdown should be specific enough that Sonnet can implement without you noticing a difference.
trvz 2 hours ago||
Maybe they could put Fable onto creating a website that doesn't use 80% of the GPU on an Apple M2.
hank2000 1 hour ago||
username does NOT check out. so confused.
alasano 4 hours ago||
I wonder if Fable 5.5 is coming this week to drown out the OpenAI dev day announcements
bastawhiz 3 hours ago|
I'm confused by the charts comparing it to Opus 5.5. It looks like slightly lower accuracy for the same cost along most comparisons. Am I reading that right?

Is it just the benchmarks? Because otherwise it suggests it's twice as chatty as Opus for a comparable output... Which kind of defeats the purpose

More comments...