Top
Best
New

Posted by D2OQZG8l5BI1S06 4 hours ago

Sonnet 5.5(www.anthropic.com)
467 points | 316 commentspage 2
robertclaus 1 hour ago|
These benchmark results keep getting more questionable without error bars.
gregwebs 3 hours ago||
This is better priced than Opus for tasks that are token heavy but not complicated. But a quick look shows that at least on some benchmarks DeepSeek performs as well and of course the cost is an order of magnitude less.

From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.

OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.

guilhermeasper 2 hours ago||
AI companies these days releasing models every week like Netflix episodes.
verdverm 2 hours ago|
the (Ai) factory must grow!
heyjstn 4 hours ago||
Have anyone tried a workflow that:

- Fable 5.1 for planning/adversarial reviewer

- Opus 5.5 for well-scoped tasks break down

- Sonnet 5.5 for these well-scoped tasks implementation

I think the blocker might be how efficient the context is compacted and sending around between these agents

robwwilliams 16 minutes ago||
Agree with afro88. Opus 5.5 as competent as Fable 5.1 on complex adversarial review of material and equations planted with errors. I still use both for a bit of variety.
afro88 3 hours ago|||
Opus 5.5 in my experience outshines Fable 5.1 anyway. May as well have Opus do plan, breakdown and review, and Sonnet implement.
chrismustcode 3 hours ago|||
You might as well use Opus for everything there.

Changing model would be cache busting spiking usage for no good reason when Opus can do it all.

Haiku 5.5 might fit well though depending on pricing.

SirMadam 3 hours ago|||
Do subagents share context? If Opus delegates to a different Sonnet window, I don't believe this busts cache?
manquer 3 hours ago|||
Context needs to pre-filled into a GPU memory in a node (usually 8xB300 or 8xH200) so there isn't any context or cache sharing between model families given their different parameter sizes, tokenizers, unlikely they are co-located in the same node.

Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.

Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]

This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.

[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it

enraged_camel 3 hours ago|||
Subagents don't share context. But that's why delegating implementation to a subagent doesn't work well except for things that are truly mechanical in nature: the subagent needs to independently reason about the task it is given, and then the output will also be reasoned about by the main agent. So you end up wasting time and tokens.
mnicky 3 hours ago|||
On the contrary, subagents save context overall, when the task is sufficiently large.

Also, my experience is that Fable 5.1 is very good at prompting/orchestrating Opus/Sonnet subagents when working on a larger task (e.g. 1-2M context window use only for the orchestrator itself).

esafak 1 hour ago|||
If you use subagents your main agent won't need to compact as often, with the loss of information that entails.
dbbk 2 hours ago|||
Using advisors doesn't break anything
Aboutplants 3 hours ago||
Do you even need Fable for much of anything now? I’m basically using it as a reviewer at the end of whatever I’m working on, and even then I’m really not finding much benefit.
nanook 3 hours ago||
Sonnet is 1/5th the price and seemingly more powerful than fable (the model that was too powerful to release). I can't make sense of this. Why would anyone use fable now? Or are the benchmarks completely pointless and one has to just try em to get a feel for what they can and can't do?
helloplanets 2 hours ago||
Wouldn't make sense to use anything below 5.5 from Anthropic at the moment. But pretty sure this is just an awkward transition phase of at most a week or two until Fable 5.5 is out.
solenoid0937 2 hours ago||
Benchmarks were always barely useful to begin with. Gotta actually try the model.
avree 4 hours ago||
Crazy bad front-end design. Site hijacks my gestures so I can't swipe back anymore, starts with a full page autoplaying video...
iAMkenough 2 hours ago|
Agreed. I’ve found that Anthropic [dot] com at least honors “reduce motion” accessibility settings, and that makes their site a bit more useable.
yapfrog 4 hours ago||
From the graph it looks like I'd rather use Opus 5.5 High than Sonnet 5.5 at all
alansaber 4 hours ago||
Always key to include the one bench where the smaller model inexplicably outperforms the larger model
mchusma 1 hour ago||
I feel like sonnet is priced too close to opus right now. If Sonnet 5.5 were half its current price it would make sense to use. At its current prices, I won't use it in applications (I would use cheaper models) and I won't use it in my subscriptions ( just use Opus instead). At least that is my initial reaction.
dom96 3 hours ago|
I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer, instead it returns to ask the user questions whether to keep going.

https://bench.killswitch-lang.org/

    Claude Sonnet 5    17.8%
    Claude Sonnet 5.5  7.4%
More comments...