Top
Best
New

Posted by D2OQZG8l5BI1S06 7 hours ago

Sonnet 5.5(www.anthropic.com)
540 points | 370 commentspage 5
pookieinc 7 hours ago|
It's interesting that in all their benchmarks, they omit Fable numbers and only focus on Opus, Sonnet, and OpenAI models. Maybe Fable is out the door?
radial_symmetry 7 hours ago||
Fable is no longer on the price/performance pareto frontier. They will probably release an updated Fable at some point that will be frontier intelligence until the next Opus.
WinstonSmith84 7 hours ago|||
"Their" benchmarks (and not just Anthropic's) look sssooooo suspicious that they would probably manage to rank Sonnet above Fable for some of their tasks which would just be next level non-sense ..
jrflo 6 hours ago|||
Models are getting more efficient far faster than they are getting more intelligent at the moment. From a marketing angle it's more impressive to focus on that, and fable would look orders of magnitude more expensive for only marginal gain, distracting from what they're trying to show here
lanthissa 7 hours ago|||
cutting edge fable is for them not you and they're not going to share the metrics until they give you access.
bpodgursky 6 hours ago||
Fable 5.5 probably drops soon so it would just be confusing.
low_tech_punk 4 hours ago||
The documentation mentions error code "frontier_llm": The request could assist the development of competing AI models.

I'm very curious how do they know what requests could assist competing AI models.

taurath 7 hours ago||
After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.
nicoburns 6 hours ago|
5 was definitely bad. 5.5 seems a lot better so far. But still not close to Fable in terms of quality.
KerrAvon 4 hours ago||
what was the problem with 5? to me, it seemed like the first Opus since 4.6 that was a real step up in intelligence without any obvious downsides
nicoburns 3 hours ago|||
It was really verbose and pedantic. I'm sure that made it more thorough. But compared to Fable (which it wasn't much cheaper than) where you could get the same rigour and more with a lot more concision, it was a tough sell. 5.5 is a lot cheaper and seems a lot better balanced.
taurath 2 hours ago|||
Read its output, and especially comments
s314 6 hours ago||
In the Artificial Analysis Intelligence Index, Claude Sonnet 5.5 is the second best model behind Opus 5.5. This however is with max effort which costs even more than Opus 5.5 max. But Sonnet 5.5 xhigh is cheaper than Opus 5.5 xigh and matches GPT 6 Astra xhigh in the benchmark.
zozbot234 6 hours ago|
> In the Artificial Analysis Intelligence Index

lol, MiMo 2.6 Pro basically matches Sonnet 5.5 high (mind you, not xhigh or max) at a far lower price point.

itishappy 5 hours ago||
Wow, I've never seen a site break chrome this badly. I get a black screen then it stops rendering the entire window, even when opened in the background.
hank2000 2 hours ago||
username does NOT check out. so confused.
sergdigon 3 hours ago||
Am I getting out of touch or is it becoming kind of confusing what model should be used when? Sure you have tons of benchmarks pareto cost/perf curves etc but at the end of the day when I have a task to give to a model it is not so clear which model and which effort I should choose ... Also benchmark numbers are often reported with max effort but by default effort is medium and based on the pareto curve on this page, Sonnet 5.5 seems more cost efficient than opus only if effort is low or medium!
mroche 5 hours ago||
Is there ever any focus on producing new Haiku models? There are a lot of use cases for quick to return models when you're limited to a single provider.
takerofnaps 7 hours ago||
Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.
__jl__ 5 hours ago|
artificialanalysis.ai benchmarks are [here](https://artificialanalysis.ai/articles/claude-sonnet-5-5). Anthropic is back at spot 1, 2, 3 and 5. Impressive even if these benchmarks are problematic in many ways.
mnicky 3 hours ago|
It's a bit misleading I think because these benchmarks are for Max level, at which Anthropic newest models use crazy amount of reasoning tokens. And we know that intelligence scales with their number.
More comments...