Top
Best
New

Posted by km144 1 hour ago

Claude Opus 5.5(www.anthropic.com)
334 points | 400 comments
sailingparrot 1 hour ago|
> Claude Opus 5.5 is our first release since we called for pacing the frontier.

Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.

mukmuk 45 minutes ago||
“Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself
DiggyJohnson 39 minutes ago|||
I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.

Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.

mpalczewski 19 minutes ago|||
The meaning isn't clear at all. So open for interpretation that it is meaningless. That's the whole fucking point. For all I know they are "pacing the frontier", or not. The fact that there's no meaning to it let's you know that it was a pointless waste of tokens and attention.
post-it 36 minutes ago||||
Is the meaning clear? Nobody would use "pacing" in this way. I only know what it means because I've seen previous press releases; if someone told me they wanted to pace the frontier I would have no idea what they mean.

I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.

sigmar 31 minutes ago|||
It makes sense to me. If you 'pace your running', you're setting the speed intentionally. The phrasing doesn't describe whether it is fast pace or a slow pace, but it describes having a goal and not just winging it.
johnisgood 29 minutes ago||||
Granted I am not a native English speaker but I have no idea what "pace the frontier" means. When I read it I just assumed "frontier" refers to "top of models" and "pacing" is that they are getting there quick.

Is this the meaning or do I have it wrong? I have not checked.

LanceH 1 minute ago|||
Doesn't it mean "restrict competitors"?
wren6991 21 minutes ago||||
It's the opposite: pacing here means "slow down" while trying to avoid the negative affect.
lxgr 14 minutes ago||||
It's actually so ambiguous that I'd sanction tabling the issue and revisiting biweekly.
ck2 11 minutes ago|||
to pace = to regulate

but without using the word "regulate" which is a negative connotation to business

but a "pacer" would be a leader of a pack which is a positive spin

it's classical business marketing language silliness

tetha 8 minutes ago||||
I'm on the fence there.

To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can "pace yourself to reach the festival by bike in about three hours to not gas out". This means to control your speed and time investment intentionally so you don't run out of energy or steam and run into leg cramps before your goal.

But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.

neo_doom 18 minutes ago||||
I suppose it depends on your life experiences. In running, someone who paces the group or a pace car is meant to keep the pack progressing at a constant, predictable speed. So in that way, it makes sense to me
arw0n 23 minutes ago||||
The meaning was immediately obvious to me as a non-native speaker, and it sounds quite poetic. Pace makes complete sense in that this is perceived as a race, and 'the frontier' is pretty much the shortest, clearest way to say 'state of the art development of AI'.
hencq 16 minutes ago||
Right, except they mean the exact opposite in this case: they're actually advocating for slowing down the pace. Hence the criticism of the language, because your interpretation would be completely valid.
fragmede 6 minutes ago||
What does the pace car in a race do? Aka safety car?
qlte 24 minutes ago||||
Yeah, when I first saw it referenced I assumed it was from something Dario wrote previously and was now disavowing, meaning "keeping up with the frontier" (i.e. racing forward from behind to match pace). Like from back when Anthropic was founded to promise they'd quickly catch up with OpenAI or something.
derac 29 minutes ago||||
In racing a pace car is a car that leads the pack and sets the pace, for instance.
browningstreet 18 minutes ago||||
It’s common terminology among runners and all kinds of racing.
lxgr 6 minutes ago||
The good old sports-to-corposlop pipeline.
squidbeak 22 minutes ago|||
> Nobody would use "pacing" in this way.

A world exists beyond your vocabulary, post it. Apparently, quite a big world.

OJFord 1 minute ago||||
> I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.

Wtf is the meaning? Means absolutely nothing to me having not seen the apparent announcement last week introducing the obscure term.

patcon 16 minutes ago||||
Agreed. This feels like pointless navel-gazing and a strange new language policing, and not something I look forward to. It's tiring and it lacks curiosity (e.g. "what's behind the affinity for the words elevated from wherever it is they come from?").

Imho people should just respond to actual ideas instead of constantly engaging in the second-order critique of how the language may or may not have been created.

It strikes me as the intellectual equivalent of "gossip" to be constantly engaging in second-order commentary on words. Of course gossip has its place and purpose, but if we seem to only let our minds live at that level, we're not moving between all the required scales of thinking that are required of this moment imho <3

lxgr 9 minutes ago|||
You must be pretty new to (not just) online discussions if you consider language policing to be a new phenomenon :)

Personally I consider it equally valid for people to publicly express annoyance with somebody's choice of words and for everybody to completely ignore that annoyance.

platinumrad 14 minutes ago|||
The words have the very practical problem of not communicating anything of substance. I guess "we want regulatory capture" didn't have quite the same ring.
DiggyJohnson 7 minutes ago||
How do they not communicate clearly? A few weeks ago they stated that they would like to regulate the pace of frontier model development/releases, and they reminded us of that in this post.
vmnb 33 minutes ago||||
people are here because they are sick of being productive
kadushka 31 minutes ago|||
We are being productive here!
vasco 22 minutes ago|||
It's compiling! Erhm... Combobulating, actually.
isoprophlex 32 minutes ago||||
You're really verbing the noun on the discourse here, belt and suspenders-style
DiggyJohnson 5 minutes ago||
what?
rythmshifter 14 minutes ago||||
I don't disagree with you, but I have to remind you

sir, this is a hacker news thread

Ar-Curunir 17 minutes ago||||
You’re acting like people are being grammar nazis when Claude and the ilk are actively making a mockery of language.
DiggyJohnson 6 minutes ago||
What? How am I doing that. Claude writing style and the discussion at hand are two entirely separate issues that I haven't conflated at all. How am I "acting like people are being grammar nazis"?
Forgeties79 23 minutes ago|||
Word choice matters
tclancy 28 minutes ago||||
It combines the elegance of LinkedIn-speak with the humbleness of desk-bound people who speak in military metaphor.
mpalczewski 18 minutes ago||
good use of AI to generate this.
nonethewiser 22 minutes ago||||
IDK I think Anthropic is plenty smarmy and weird itself. Sounds like they wrote it.
topbanana 28 minutes ago||||
You're right to call that out
marton78 28 minutes ago||||
Sounds like "flatten the curve" and "the hammer and the dance", both coined well before AI.
tclancy 27 minutes ago|||
I learned the former from Waylon Jennings and the Dukes of Hazzard.
pvab3 26 minutes ago|||
except with a good dose of EA and Star Trek added in
Dumblydorr 34 minutes ago||||
What specifically is smarmy and weird? Sounds like your own hot take with zero analysis.

They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.

Do you have a better proposed phrase?

Rebelgecko 14 minutes ago|||
I thought pacing is just walking back and forth? So pacing the frontier would be like staying in the same place instead of making progress
plaidfuji 14 minutes ago|||
“… since we called for a slowdown in AI research”

But stating it plainly like this would make the contradiction too obvious.

gradus_ad 34 minutes ago|||
Agreed when I first heard the phrase it sounded odd. Maybe they thought it subtly conveyed they would be setting the pace... But again this is something AI would come up with in its awkwardly post hoc sort of way.

Though tbf corporate-speak and AI-slop are both insufferable in similar ways...

dmazin 58 minutes ago|||
It seems like they are. I mean, this is similar in performance to Fable (ish). It seems like more focus on making existing capabilities more accessible.
sailingparrot 53 minutes ago|||
Fable 5.1 came out just 21 days ago. Only 3 weeks! And this is 20% relative improvement on terminal bench vs Fable 5.1 at less than half the price, and more human sounding output. does not feel paced to me tbh.
davrosthedalek 41 minutes ago|||
Well, I guess it's "fast paced".
jr3592 49 minutes ago||||
What exactly is "paced" in this context?
sailingparrot 47 minutes ago||
It’s the famous “flattening the curve” from COVID. But for LLMs. This release is not flattening anything.
sidrag22 6 minutes ago|||
This is a preexisting model being optimized. Its absolutely not some unexpected release after that blog post. I won't defend that blog post, but saying THIS release is proof they don't mean they are slowing down is just incorrect, this is a prime example of what i consider horizontal improvements

Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.

The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.

jr3592 32 minutes ago|||
I guess I understand why we'd want to flatten a COVID curve, but why do people want to flatten the LLM development curve? Don't we want the opposite? Isn't the goal AGI?
sailingparrot 24 minutes ago|||
There is a difference between wanting AGI (which not everyone does), and wanting it as fast as possible no matter the side effects and potential for vast harm. Homo sapiens is 300k years old, maybe it’s ok to delay AGI by like… 1 year if it meaningfully improve our ability to align the model?
lantry 30 minutes ago||||
Well, there's a tension because, depending on who you ask, AGI is how you cure cancer and achieve utopia, but also how you kill all life on earth and turn the solar system into paperclips
recursive 15 minutes ago|||
I think the goal is different from what "we" want anyway.
dmix 39 minutes ago|||
Fable 5.1 wasn't that much different than Fable 5 though.
cab648bec139cc 54 minutes ago||||
Do you guys still believe any of their lies? You are getting trolled for years by now and yet you still believe what they tell you?
reasonableklout 2 minutes ago|||
I mean they are literally getting sued since 3 days ago for trying to coordinate a slowdown, there is a very clear reason why they cannot effectively self-regulate.
supern0va 48 minutes ago||||
That's a great point, five minute old account.
sleazebreeze 53 minutes ago|||
What do you think is happening?
re-thc 50 minutes ago|||
IPO soon
cab648bec139cc 49 minutes ago|||
I have no idea what is happening. I just know that programmers will not be replaced in 6-18 months (tm).
anthonyrstevens 36 minutes ago|||
Who said that, why do you take their word as the literal truth, and most importantly, what does this have to do with a focused discussion of Opus 5.5?
felixgallo 4 minutes ago||
the prompt said that, so of course the anti-anthropic bot took it as literal truth.
meowface 25 minutes ago||||
Dario never said they would be. Just that more and more code will be produced by LLMs. All his predictions were in fact pretty much right in terms of months and percentages, give or take small margins.
0xbadcafebee 42 minutes ago|||
Maybe not replaced exactly but they won't be manually typing out lines of code anymore. I haven't written a line of code in like 6 months. I review PRs, write prompts and tickets, check CI output, and get frustrated when the magical code machine stops working or I run over token budget
the_gipsy 56 minutes ago|||
Occam's razor: they couldn't make any more substantial improvements.
dgellow 38 minutes ago|||
A razor is a philosophical tool to help decide between options, in the case of Occam it’s a way to decide for something in a situation where multiple options have more or less the same level of plausibility to en your current knowledge. It’s a heuristic to make a “cut”. What are you shaving off?
someothherguyy 54 minutes ago||||
> Occam's razor: they couldn't make any more substantial improvements.

doesn't sound like a razor at all

bpodgursky 55 minutes ago|||
Everyone knows both labs have internal models which outperform the frontier. All releases are to match market parity and demand for spend, the rest of the compute is used for training. It's not worth arguing about this.
nextaccountic 41 minutes ago||
Maybe their internal edge dried up in the last months
lukewarm707 52 minutes ago|||
the only thing they are pacing is what models the permanent underclass are allowed to have in life.

that, they fully intend to 'pace'. hiring accenture is a good sign they need some justification theatre and a fall guy for this decision.

drnick1 33 minutes ago|||
With AI tools it's easier than ever to create a business, do research, or build stuff. That's an opportunity for the "underclass," not a curse.
azan_ 40 minutes ago||||
Either AI will capture so much value that there will be permanent underclass (and in this case it's extremely capable and extremely dangerous and should be heavily regulated) or it won't be capable enough to displace people into permanent underclass.
user3939382 44 minutes ago|||
Or they’re running into a steep diminishing return slope on R&D vs performance and are using stewardship as a cover.
Iolaum 44 minutes ago|||
They are advertising the regulations they want to enforce in the following sentence, which makes their intentions explicit (ie apply those things made to suit us to our competitors).
tencentshill 20 minutes ago|||
What an amazing excuse for lower than expected performance! Our models are slow because we're so ethical.
felixgallo 3 minutes ago||
in what way is beating every other frontier model with their own second-tier model, 'lower than expected performance'? Please be specific.
qgin 14 minutes ago|||
This IS pacing. Nobody said pacing would mean slow.

Pacing is very explicitly about RSI and similar training methods that will accelerate progress beyond our ability to comprehend it.

dspillett 20 minutes ago|||
When they talk about pacing, they are referring to their dangerous competitors, particularly those evil open-sores and Chinese ones, not their lovely safe models because you can trust them to look after your interests.

What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.

kadushka 51 minutes ago|||
This makes perfect sense. There are no real improvements anymore (just benchmaxxing), and they explain it by "pacing the frontier".
heyjstn 24 minutes ago|||
Others must slow down, but not us.
scottyah 54 minutes ago|||
Seems like a bigger focus on efficiency (both cost and speed) and the "tone" of Claude vs benchmarkmaxxing
dr0idattack 47 minutes ago|||
a 1 minute mile pace
AtlasBarfed 22 minutes ago|||
If they were really about putting brakes on these, they would simply make these things non-agentic.

Simply make them something that derives a text response from its training data.

CodingJeebus 54 minutes ago|||
It's laughable at this point. It feels like they're drumming up all this fear about imminent AI threats to emphasize the need to slow down, when in reality, the model progress seems already to be slowing down and has shifted to compute allocation (i.e. "how much compute do you want to throw at this prompt?"). All while continuing to tout benchmark records with each new release.
jr3592 25 minutes ago||
This. I swear the fear mongering is all about investor signaling and regulatory capture. It's so disgusting that anyone believes it.

The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.

BatmansMom 56 minutes ago||
kinda disingenuous. They include a whole section on pacing later on
sailingparrot 49 minutes ago||
You mean the section where they tell us this model is not affected by pacing because “they understand it well” and they will share more details on pacing later? Yea not very convinced by this effort.
GodelNumbering 1 hour ago||
Finally that price drop

   Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
   Cache reads              $0.20              $0.50
   Input tokens             $4                 $5
   Output tokens            $20                $25
   Cache writes             $5                 $6.25

Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.

If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor

AJ007 43 minutes ago||
It is only a price drop if price * tokens used is less
mcintyre1994 34 minutes ago||
They're claiming a drop in token use too, and that it nets to 40% cheaper.
drbscl 30 minutes ago|||
Unfortunately, they're full of it https://artificialanalysis.ai/models/claude-opus-5-5#token-u...

It does work out to be a similar cost per task though

jsnell 22 minutes ago|||
You should probably look at the cost/score graph by effort level instead:

https://artificialanalysis.ai/models/claude-opus-5-5#intelli...

It is most of the pareto frontier.

drbscl 17 minutes ago||
Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort
naasking 20 minutes ago|||
I don't think so, I typically use Opus 5 on High, and 5.5 scores lower on token use:

https://artificialanalysis.ai/models/claude-opus-5-5?models=...

make3 15 minutes ago|||
parent means that they could get more client / a larger part of the market, which would lead to more income (more tokens) despite lower marginal prices
Shekelphile 6 minutes ago|||
Footnote on their pricing page says:

> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.

If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.

coffeebeqn 46 minutes ago|||
We haven’t been able to use opus as much as we’d want because it’s been too expensive for general use, price drop is good so I can stop juggling different models and just use this daily unless it has some weird new issues
chrisweekly 28 minutes ago|||
Price per task (not per token) is what really matters.
cute_boi 26 minutes ago||
Agree. But similar to how ISP use 200 mbps (bits) instead of 25 MBPS(bytes), i think this trend isn't going away.
blfr 47 minutes ago|||
People are paying for Opus 5? Not just burning down tokens left after they enjoyed Fable on the sub? Amazing.
rapfaria 43 minutes ago|||
My workplace doesn't even offer Fable. And on the sub, I've had a hard time understanding Opus 5, but Fable can deal with it with subagents.

If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate

blfr 39 minutes ago||
Telling Fable to delegate is agentic development. At least I thought so until reading your comment.
neuronexmachina 27 minutes ago||||
Enterprise and most Team accounts use API pricing, they don't have an included-usage quota.
herpdyderp 27 minutes ago||||
When you need to disable data retention, you cannot use subscription plans.
ascorbic 26 minutes ago|||
Enterprise, and APIs
alvis 1 hour ago|||
60% cache read is cool, but subscription only get 25% more according to Cat. I'm confused
bayesianbot 44 minutes ago|||
I think gpt 5.6 family also dropped pricing but didn't give any more usage for the subscriptions. Maybe it's a way to silently lower the value given to subscriptions while keeping API pricing competitive
weiran 50 minutes ago||||
25% more usage sounds about right given the other token costs are down about 20%? I don't think cache read is a big portion of the overall cost.
Espressosaurus 44 minutes ago|||
Anything with long context quickly gets dominated by cache reads. Especially for interactive sessions I’ve got cache read % between 95% and 98%.
hedgehog 13 minutes ago||
In my mix it's usually 98% or 99% at which point Fable 5.1 was pretty close to the same cost as Opus 5 due to the cheaper cached read. I've seen similar numbers for other people with long-running tasks running experiment loops and than sort of thing.
re-thc 49 minutes ago|||
> I don't think cache read is a big portion of the overall cost.

For long running tasks it is. That's what made Deepseek so cheap.

vardalab 4 minutes ago||
Yeah, flash models, DeepSeek, MiMo, GLM, I love those things. For simple tasks like a daily routine shit, just setting up stuff and then doing the hard stuff in Claude/Codex, that's a reasonable approach for someone like me, a "gentleman code farmer", lol. And even lower tier stuff, I have the local models taking care of. Now that Jev is out I can finally have a true AI sysadmins managing my "cloud in the basement" homelab at the cost of electricity, which is not cheap btw
liudaisuda 51 minutes ago|||
source link please?
rahimnathwani 39 minutes ago||
"Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that."
mcintyre1994 32 minutes ago||
> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.

I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I've don't think Astra is a better model, but it's the first OpenAI one that seemed good enough for me. Definitely keen to try Opus 5.5 and see if this claim is real.

epicepicurean 1 minute ago||
Much better than Opus 5. prompt:

> hi, can you explain how the scheduler works. keep it brief, but include important correctness details

some excerpts:

>Flow: 1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT. 2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot. 3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.

> - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.

> - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.

> - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.

All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.

derangedHorse 21 minutes ago|||
As someone who uses both, Astra was 100% the better model. I have yet to give 5.5 a spin so maybe that’ll be the new top contender.
jaflo 24 minutes ago|||
I did the same switch (that reason along with the newer models seeming more "lazy" and needing constant prodding to finish long-horizon tasks) but my issue with ChatGPT/Codex now is that it too roundabout and doesn't get to the point. I tried adding instructions and using the personalization settings to make it more efficient but haven't seen much change. Claude seemed to follow settings more closely. Has anyone had any success to make ChatGPT more succinct?
algoth1 25 minutes ago|||
Please update with your feedback
sha-3 14 minutes ago||
I haven't heard it say "load-bearing" yet (I've used it for 30 minutes now), so that's a start.
Trasmatta 26 minutes ago||
Opus 5 has made me question my sanity on a daily basis, especially as all my coworkers started lobbing Opus 5 slop grenades everywhere. It had the worst and most infuriating writing style I've ever seen.

I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.

One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.

nonethewiser 3 minutes ago|||
I really wonder how it converged on its style. It's pretty unique and terrible. It's not like it's just mimicking something or it was purposefully design to be that way. I mean the reason may be diffuse and uninteresting... just the result of a lot of factors and lack of control over the writing style probably.

But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.

LtdJorge 13 minutes ago||||
Yes, it made me want to vomit. If the new Fable only changed the writing style to just sound like a human, same performance for everything else, I'd be pretty happy.
Aperocky 19 minutes ago|||
It's not X, it's Y, not A, not B, not C, and he haven't even woken up yet! Here's the catch, the detail is in the devils and the twist is that it's designed!
simonw 24 minutes ago||
Here are pelicans for thinking levels low, medium, high, and xhigh: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.

I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!

Max started its thinking trace like this:

> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.

So that failed attempt on max cost me $2.56.

I ran this using my llm-anthropic plugin:

  uv tool install llm
  llm install llm-anthropic --upgrade
  llm keys set anthropic
  # paste key here

  llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"

  # Then to save the markdown logs
  llm logs -cu > logs-with-usage.md
MikhailTal 23 minutes ago||
> This is a classic test request

Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?

Brendinooo 17 minutes ago|||
Plenty of times I’ve seen a model say “it’s a classic X” despite not being a classic anything. Might just recognize it’s a test in general, or it might just be a tic.
copperx 4 minutes ago||
[delayed]
MaxikCZ 18 minutes ago||||
Dont conflate "I know this is test case" with it being trained on it.

But its safe to say that pelicans on bicycles are disproportionally huge part of their training data

simonw 21 minutes ago||||
It's the model admitting that it has heard of the test. It's been around for a couple of years now so I'd be surprised if it hadn't.

Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.

zamadatix 19 minutes ago||||
I think people just like to see the drawings at this point.
segbrk 16 minutes ago||||
Not really. Of course it has pelican benchmarks in its training data. It likely has every article linked on HN in its training data. But that doesn't mean it was "trained on" the benchmark, as in specifically fine-tuned to make a better pelican. It just "knows" that the request is a benchmark.
FergusArgyll 21 minutes ago|||
It has read the internet. That doesn't mean it was literally RL'ed for this
ealready_value 18 minutes ago|||
I agree, not a huge difference here. They eyes and ... hat? on high are out of place so I'd argue that's the worst one, but it takes xhigh before we get legs and bike ordering correct.
Kurtz79 22 minutes ago|||
Heh. Pelican-benchmaxxing is real.
cainxinth 19 minutes ago|||
I guess that means you are officially the creator of a "classic" LLM test. Congrats!
inshard 15 minutes ago|||
Not as good as Astra or Fable 5.1 on this test as far as I can see. I wonder if any benchmark exists for artistic taste, visual sophistication etc. I think your Pelican test does touch on these aspects of a model and is useful for developers trying to build rich digital experiences (includes games, interactive websites and apps). These benchmarks are subjective so it may not be easily established and will have polarized reactions before it gains legitimacy. May even need human judgement layers adding to the cost of running it.
make3 12 minutes ago||
this benchmark is useless and should die
copperx 2 minutes ago||
[delayed]
techjamie 1 hour ago||
With the performance gains they're claiming, I wonder if they implemented the Casual Encoder-Decoder technology from DeepSeek 4.1's paper.

I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.

How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...

stri8ted 7 minutes ago||
This model was likely trained months before deepseek released their paper.
ryangg 36 minutes ago||
Getting a 403 on that link. Mind checking it once?
peri-cl 21 minutes ago|||
https://web.archive.org/web/20260922172456/https://miraflow....

tired: AI startup attempting to publish a webpage

wired: a nonprofit founded in 1996

potwinkle 33 minutes ago||||
I'm able to access it on my laptop at home. Maybe a misconfigured bot protection rule, try a different user-agent or IP?
zatkin 29 minutes ago|||
It's working for me (based out of California).
ApolloFortyNine 1 hour ago||
>Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.

peri-cl 1 minute ago||
I love the contrast with the open-source MiMo release yesterday, which put research chemistry (metal-organic frameworks stuff) front and center in the release notes.

https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials...

sys32768 3 minutes ago|||
Fable and now Opus 5.5 won't answer my college student's prompt about Alzheimer's and immune response.

ChatGPT 6 Pro answered it without issue.

blfr 40 minutes ago|||
Fable 5.1 addressed an entire security advisory I had that Fable 5 and Opus 5 refused. I think they loosened the leash a little.
arw0n 18 minutes ago|||
It has far less false positives now, and generally accepts defensive requests. When it comes to offense, you can actually ask about certain types of vulnerabilities if you phrase things carefully, but it will block hard if it is about exploits.
cute_boi 24 minutes ago|||
If they don't loosen, people will choose Astra or Chinese model.

Giving moral lecture is different than reality i guess.

prettyblocks 1 hour ago|||
They're pushing their customers to their own competition by doing this.
Espressosaurus 37 minutes ago||
It’s not like ChatGPT isn’t doing similar. I’ve been hit by cybersecurity strikes before while working on an internal codebase that I had to appeal. Anthropic hasn’t done that to me yet. ChatGPT also regularly does that “thinking for a long time while we check if your chat is rule breaking” thing a lot for me when doing model identification without even interacting with external codebases or services.

The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.

Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.

raesene9 23 minutes ago||
For some cybersecurity tasks, the Chinese models are already good enough, things like PoC development or things like exploiting mis-configurations.

Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.

KeplerBoy 29 minutes ago|||
Anything else would be inconsistent, wouldn't it?
searine 51 minutes ago|||
Great. Claude is basically useless for bioinformatics now.
unglaublich 47 minutes ago||
Opus is useless; Mythos access will be granted to companies that are friendly to the government, so the government gets more control over business.
bushido 37 minutes ago||
One of my favorite things about their safeguards is their own model will utter something which it does not like and then I'll need to reset the conversation.

The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.

ACCount39 3 minutes ago||
You ask it about some thing, then you see it tangent into "things like that are sometimes used in biomedical applications like-" and then it just shoots itself in the head. Wonderful.

That kind of bullshit was the old Opus filters too.

If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.

m4tthumphrey 1 hour ago||
Just post the bloody content. This UI/scrolling thing is horrific.
amluto 40 minutes ago||
Claude Opus 5.6 should have a new "UX safety" feature that requires annually-renewed preauthorization to generate webpages that hijack scrolling :)
swader999 53 minutes ago|||
I told my team to smack me upside the head if I ever try to ship something so daft as that.
gruez 1 hour ago|||
???

It's just a standard hero image + text for me, with no scrolling effects.

edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.

KyleTheDev 59 minutes ago|||
If you're at the top of the screen, at least in Chrome 153.0.8010.37, it has a little interactive bit. You have to scroll through the images in order to be dropped at the actual web page, at which point the images go back to being a regular part of the page.

I agree that it's sort of stupid, not a fan.

ealready_value 22 minutes ago||
It's less than OpenAI did for Astra, but that was my first encounter opening it and my first thought was that they decided they liked Astra's hero/scrolling animation. I'm pleased to see they didn't make the entire page that like OpenAI did, but I'm expecting to encounter this pattern more often on these announcements now.
EricBurnett 1 hour ago||||
Two posts were merged; this comment was for the blog post with an intro animation thing.
mbreese 56 minutes ago||||
On mobile at least, you have to scroll to get the TOC to appear. Then keep scrolling to actually move off from the hero to see the text.

For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.

thejazzman 1 hour ago||||
then you're getting served a different website
giancarlostoro 58 minutes ago||||
On mobile its different.
iAMkenough 57 minutes ago|||
Figured it out: you have "reduce motion" enabled in your device's accesibility settings.

Everyone that doesn't gets served some animated bullshit.

gruez 8 minutes ago||
>Figured it out: you have "reduce motion" enabled in your device's accesibility settings.

Yep, you're right. I tried on my phone and got the scroll through image.

thebitguru 55 minutes ago|||
Totally! So unnecessary and annoying.
halyconWays 55 minutes ago|||
I call it scrollslop
josefresco 38 minutes ago||
Hijacking the scroll wheel has existing long before "AI". Many "high end design" websites that want to "tell a story" get woo'd into thinking it's a good idea. It's terrible, and feels like your scroll wheel is stuck in quicksand.
halyconWays 17 minutes ago||
Those sites are also scrollslop. "Slop," as a term, is independent of AI
iAMkenough 55 minutes ago|||
Turn on "reduce motion" in your accessibility settings and you get served a sane version.
dionian 59 minutes ago||
and hijacking back/forward
2001zhaozhao 2 minutes ago||
It's great that we are finally getting bankable rate limit resets for subscription users. According to another comment here they apparently last a month.

I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.

This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.

sharkjacobs 1 hour ago||
> “Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it

God I hope so

lgessler 39 minutes ago||
I thought about taking a shot every time Opus 5 said "load bearing", "bites", "teeth" (real oral fixation it had), "real {concern,issue,problem,...}" and realized I'd be dead of acute alcohol poisoning by lunch if I did so.
drbscl 29 minutes ago|||
So did I. Unfortunately it's even more verbose according to https://artificialanalysis.ai/models/claude-opus-5-5#token-u...
Trasmatta 23 minutes ago||
The problem with 5 wasn't just the verbosity, but its insane way of communicating. It had this bizarre circuitous sentence structure that always buried the lede, and always tried to be faux profound. I'm okay with verbosity if it's actually readable.
fastball 35 minutes ago|||
It hasn't just fixed it, it has introduced a new paradigm in anti-obscurity.
mikeocool 34 minutes ago|||
That's the load-bearing seam in this blog post.
boc 36 minutes ago|||
So far in the past 20 minutes it sounds much better in my sessions. Way better than 5.0 so far.
neilellis 29 minutes ago|||
'frontier models' - seriously, it was you and only you!
unddoch 19 minutes ago|||
It is hilarious to me that in the examples they show side by side Opus 5.5 still uses 4 times more words than it needs to use. IME, if you eyeball how many words the thing they're trying to say actually needs, and tell them to use only this many words, they become excellent communicators. I assume something about Anthropic's grader for writing just really wants to tick all its tidy tiny boxes of information the models need to cite. It's terrible.
kantahayashi 52 minutes ago|||
The improvement in writing sounds great! I want OpenAI to follow it. Writing in recent models is a disaster.
bushido 47 minutes ago||
Install the simple English skill. OpenAI follows that really, really well.

https://github.com/AminBlg/SimpleEnglish

mavamaarten 48 minutes ago||
That's literally all I'm hoping for. Is it an insufferable cunt and does it write awful text, or is it nice to work with?
joshstrange 1 hour ago|
> It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.

Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!

bayesianbot 1 hour ago|
Wow those cache reads are quite reasonable - I think that's equal to 5.6 Terra. I might have to try Claude again after years of being priced out of it
More comments...