Top
Best
New

Posted by alvis 6 hours ago

Claude Opus 5(www.anthropic.com)
https://www.anthropic.com/claude-opus-5-system-card
1108 points | 590 commentspage 4
petilon 4 hours ago|
The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
einsteinx2 4 hours ago||
Fable is better than Opus which is better than Sonnet which is better than Haiku. They’re basically just sizes.

Though it gets even more confusing because they also have effort levels so it’s not really possible to call one fast and one slow since Fable on Medium will be faster than Opus on Max.

I agree it’s confusing, and now OpenAI is following Anthropic’s lead with their new naming (Sol, Terra, Luna).

bonoboTP 4 hours ago|||
It's really not all that confusing. It takes 5 minutes to understand. Optimizing for absolutely no effort needed is silly. It's a thing, a topic, a skill, a domain. You have to get a little bit familiar with the terms in order to use it. Everything works like that. It's not that hard. The learning curve is very graceful. You can literally just start by asking any chatbot what the names mean. It's that easy.

A similar complaint was valid years ago when OpenAI had GPT-4o, o1, o3 (but no o2), o4-mini-high, GPT-4, and GPT-4.1 and GPT-3.5 etc.

einsteinx2 3 hours ago||
I am familiar with the terms, but I also can see how it can be confusing for a lot of people.

Arguably the complaint was more valid for those older GPT models you mentioned.

bonoboTP 2 hours ago||
Ok, the boring way would be a subset of XXS, XS, S, M, L, XL, XXL like clothes sizes. But it loses some marketing appeal and a quirky touch of personality that companies like.

Some models like ViTs use something similar but then introduce words with no unambiguous order, like Small, Medium/Base, Large but then I always forget if Huge or Giant is larger.

paxys 4 hours ago|||
But according to their benchmarks Opus 5 outscores Fable 5 on basically everything. So which one is “better”?
einsteinx2 3 hours ago|||
Maybe more accurately I should have said “larger”. Fable has the most parameters, Haiku has the fewest.

Also fwiw I’ve never found LLM benchmarks to match reality based on my own usage, not for the large frontier models or smaller open weight models so who knows if Opus is actually better than Fable (I doubt it).

tackta 1 hour ago|||
I think the problem is that Fable 5 is probably a bit outdated right now.

Fable 5.1 or whatever they go with will be the stronger version vs Opus 5.

From about 2 hours of Opus 5 use , I would say it is quite impressive.

hk__2 4 hours ago|||
An opus is longer than a sonnet, which is longer than a haiku. Hence Opus > Sonnet > Haiku.

> Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.

This is not possible: Standard (Free) / Pro / Max are plan names. Fast is a mode.

taybin 3 hours ago||
How long is a fable though?
FergusArgyll 1 hour ago||
https://en.wikipedia.org/wiki/Aesop's_Fables#Select_fables

Enjoy

abalaji 4 hours ago||
Opus is better than Sonnet -- an Opus is longer than a Sonnet
bonoboTP 1 hour ago|||
An opus is just short for "magnum opus", and it's a different type of label than a sonnet. A sonnet is a very particular kind of poem, while "opus" basically just means an important work. It can be short or long, and has no format requirements like sonnet (or haiku does).

And fables are not particularly long actually.

petilon 4 hours ago|||
> an Opus is longer than a Sonnet

And people know this? I didn't. I am not into music or poetry so these are not terms I am familiar with.

MostlyStable 3 hours ago||
I guess you are one of today's lucky 10,000 [0]

[0] https://xkcd.com/1053/

Mossly 2 hours ago|||
This inspired me to check lol. Brysbaert et al. (2019) collected word prevalence norms (the share of people who report knowing each word) for ~62K English lemmas from ~220K participants.

fable: 99/100 sonnet: 97/100 haiku: 91/100 opus: 89/100

So while these terms are almost universally known, opus is indeed the least known of the four. And I guess this only measures whether a person knows a word, not whether they know an opus is longer than a sonnet! Personally I only inferred that based on the related term 'magnum opus.'

josefresco 2 hours ago|||
Not as lucky as the guy seeing the Mentos/Soda trick for the first time!
hrpnk 3 hours ago||
The breaking changes vs. Opus 4.8 are interesting [1]

1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.

2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.

[1] https://platform.claude.com/docs/en/about-claude/models/migr...

slymax 2 hours ago|
on claude.ai it's no longer possible to disable thinking at all for Opus 5
ddxv 5 hours ago||
"Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation."

Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.

Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.

layer8 5 hours ago||
“Proportionally”? In proportion to what?
ReptileMan 5 hours ago|||
In one chat - can you disassmble x?

In the next - please scan this totally mine code for vulnerabilities

zb3 5 hours ago||
It will probably refuse to work on source code written by me by hand, because it might think it was obfuscated/decompiled..
redsocksfan45 3 hours ago||
[dead]
alasano 5 hours ago||
Half the price of Fable 5 and useable with 100% of your subscription means roughly 4x the usage using Opus 5, presuming similar token use for solving problems.

Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.

mchusma 5 hours ago|
There is a expiring soon 50% boost to your usage limits, so I think its 2.7x not 4x what you are seeing right now. I think, but its convoluted :)
tekacs 5 hours ago||
Something fun: on our AWS Bedrock console right now, there's a 'NEW' model called 'anthropic.honey'. Wonder if that's the codename just for this one or in general?
lucamark 5 hours ago||
But why GPT 5.6 Sol is so behind on the benchmarks? In real-world projects, it is the best frontier model to me in terms of accuracy, speed and consistency. It can just be compared to Fable 5, but I prefer GPT 5.6 Sol because of inference speed.

I've never trusted on model cards though. I'm sorry.

dbbk 4 hours ago|
Another benefit is that fast mode can be used on subscription, but Anthropic's won't
lucamark 4 hours ago||
Exactly! And they should also release the new inference engine in this month. Anyway, I am curious to try Opus 5, considering that previous versions (e.g., 4.8) were disappointing
modeless 5 hours ago||
Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
oh_no 5 hours ago||
seeing a jump this big is not a great sign for the continuing value of a benchmark
modeless 5 hours ago||
It will continue to be valuable as a cost and speed benchmark long after it is saturated at the high end. And they are already working on ARC-AGI 4 and thinking about going even farther.
dominotw 5 hours ago||
> I continue to believe ARC-AGI measures something different

why is that? its now being benchmaxxed too

trunnell 3 hours ago||
The chaos appears to be tamed for now.

From the system card [1]:

  The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
  If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...
williamstein 5 hours ago||
> This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.

Annoyingly, this is a concrete argument that open source software may be easier to attack.

irthomasthomas 4 hours ago|
Changelog - fixed issue where model acts like qwen when prompted in chinese
More comments...