Top
Best
New

Posted by sfkgtbor 4 hours ago

Claude Haiku 5.5(www.anthropic.com)
488 points | 229 commentspage 2
garo-pro 3 hours ago|
> Claude Haiku 5.5 is our fastest model to date at each model’s standard speed, although it runs less quickly than our Opus models in Fast Mode.

Opus 5.5 runs 117 tps average on Openrouter, so it must be at least 10-20 tps slower for them to mention. IDK why they mention this as it does not help for marketing though. https://openrouter.ai/anthropic/claude-opus-5.5

jstummbillig 3 hours ago||
Maybe they think it's of interest.
fred_dawg 3 hours ago||
Is that page showing Opus TPS stats in fast mode? IIRC fast mode is 2.5x speed, so that would be 293 TPS, no?
yorwba 2 hours ago||
117 tps is the fast one, regular speed is 69 tps.
TheAmazingRace 4 hours ago||
I wonder if we have an AI LLM equivalent to Moore's Law. Like how often do we expect improvement in this technology and with what timing?
onlyrealcuzzo 4 hours ago||
Yes -> every 18 months they've gotten 90% more efficient for the same level of quality for about 5 years. There's little sign that trend is slowing. If anything, there's reason to believe that System 1 models (plus potentially 1-2-3 workflows) may increase that over the next 3-5 years.

You'll know when the trend stops -> when the intelligence differential between smaller models like 7B starts to grow instead of shrink from 32B models -> that means 7B is getting about as smart as it can get. Then, 32B will follow next, then 70B, etc etc.

We haven't yet seen that at any size AFAIK.

thefourthchime 3 hours ago||
Andrej Karpathy said once that he expects superintelligence could fit in 1 billion parameters.
onlyrealcuzzo 3 hours ago||
Super intelligence that doesn't have to deal with the real world, maybe.

I wouldn't be surprised if less than 1B param equivalent of our brain deals with solving math and writing computer programs and physics and all the things we tend to associate with "intelligence" - especially if you ultra optimized for that, I doubt our brain works like that.

Dealing with the real world, I highly highly doubt it.

Gigachad 44 minutes ago|||
It would be interesting if running ends up being a more complex task than advanced math. And our brains are just 95% allocated to dealing with the real world.
jstummbillig 3 hours ago|||
How about if we get away from written text as the input, to something more fundamental, that then also is able to produce text (among other things)?

Given that humans learn to talk while having encountered a measly number of word instances, and, given enough time, we should always be able to improve on the lottery that is biology, it does seems fairly likely.

istjohn 3 hours ago|||
According to Epoch AI:

> The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. [0]

0. https://epoch.ai/publications/the-plunging-price-of-thought

FooBarWidget 3 hours ago||
Then why are AI plans still so super expensive, and AI spending going through the roof, while all the subsidies are ending?
stephbook 1 hour ago|||
https://en.wikipedia.org/wiki/Jevons_paradox

AI gets cheaper, people use it everywhere. Google searches, for example. Now we want to crack math problems and spend weeks with unreleased models.

If you used GPT-2, it'd be incredibly cheap. You basically can't use it for anything and it's simple to serve.

f6v 34 minutes ago||||
Reddit is full of people complaining how they burn their 200$ sub in half an hour by starting ten Max sub agents. That’s to say, many people just don’t know what they’re doing.
adgjlsfhk1 3 hours ago||||
The cost per fixed level of intelligence is dropping, but we're also getting dramatically more intelligent models.
jrflo 2 hours ago||||
Because models are only getting better at a rate of 10% per year, people always want the best quality possible. You can get SotA performance from a year ago for a fraction of the cost, but why would you use Opus 4.5 when you can use Opus 5.5?
jstummbillig 3 hours ago||||
Because it's increasingly useful and the thing you are substituting (human time) is much more expensive.
srdjanr 2 hours ago||||
Apart from what others said about using more intelligent models instead of cheaper ones, token usage is also increasing a lot. Classic Jevons paradox
teaearlgraycold 3 hours ago||||
At least for me the Claude plans seem like an incredible deal and I never hit my limit.
bravetraveler 4 hours ago|||
I've heard tell about 100% of certain types of work being ended in batches of six months. For years. Truthfully, I'm skeptical, but accuracy wasn't prioritized.
ChaseRensberger 4 hours ago|||
reminds me of this blog post: https://campedersen.com/singularity
kator 49 minutes ago||
Whew, at least I won't have to hand-code solutions to the 2K38 problem!
himata4113 4 hours ago|||
double the information density every 2 days?

serious bit: if you think about how these smaller models work, at the end of the day it seems that they are now capable of forgetting useless information because they're able to derive it in reasoning allowing models to become smaller at the cost of requiring more reasoning tokens to solve a task.

qeternity 4 hours ago||
Knowledge will be shifted to systems like n-gram augmentation which are relatively cheap and will not compete with reasoning capabilities for weight saturation.
dyauspitr 4 hours ago||
Hopefully enough runway for an existing model to train the next to be better than itself with absolutely no human intervention.
djoldman 2 hours ago||
https://www.anthropic.com/claude-haiku-5-5#further-updates

This section makes the reader think: why would I not pick Sonnet 5.5 instead of Haiku 5.5?

tpoacher 4 hours ago||
Good to see Anthropic back alternative OSes.
MisterMunchkin 2 hours ago||
> we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200

They’re definitely planning to make the subscriptions API based so they can charge you full price.

Topfi 3 hours ago||
131tok/s P50 according to OpenRouter currently, though might move up or down over the coming days. If it sticks at that speed, roughly twice the throughput of Luna and far lower latency (up to 2sec depending on provider) is impressive, though the 5x price increase beyond 100k is painful.

Was a big fan of Haiku 4.5, though understand why for most Sonnet was the far better option back then.

waximabbax 2 hours ago||
Alright its still little early since there is not enough independent testing but this looks very promising and I wasn't expecting anthropic to beat GPT-6 Luna especially at the same price. Haiku 5.5 beats Luna on every shared benchmark Anthropic published, particularly computer use and agentic coding.
TomGarden 4 hours ago||
From these selected benchmarks, it looks like it smokes Luna capability-wise. Excited to put it through its paces
swalsh 3 hours ago||
Top of the page in 17 minutes? Now I know what y'all do while your agents are working.
skeledrew 2 hours ago|
It's a brave new world... of idleness!
dangoodmanUT 3 hours ago|
> Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API.

This is kind of nuts

djeastm 2 hours ago|
Is it realistic or cynical for me to assume this is to wean developers off the heavily subsidized subscriptions? Presumably it's using similar compute.
copperx 1 hour ago||
Anthropic was ignoring the usage of third party harnesses. Not anymore.
More comments...