Top
Best
New

Posted by alvis 5 hours ago

Claude Opus 5(www.anthropic.com)
https://www.anthropic.com/claude-opus-5-system-card
1015 points | 538 commentspage 3
MasterScrat 58 minutes ago|
Damn the pelican guy can’t get no sleep
pyridines 4 hours ago||
The wording in this post seems much more... restrained? than usual. Maybe Anthropic is afraid of exaggerating the capabilities and consequences of their new models to avoid government scrutiny and sanctions.

> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities

I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.

MallocVoidstar 4 hours ago|
Opus 4.8 was intentionally nerfed and that was before the government took action against Fable
cebert 5 hours ago||
I am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.
user43928 4 hours ago||
Fable 5 is assumed to be a larger model.

It seems plausible to me that RL improvements allowed Anthropic to improve on Opus 4.8, similar to how OpenAI substantially improved upon GPT 5.5 with 5.6 Sol.

Fable 5.1 and GPT-6 are rumored to launch in August, presumably bringing those improvements to the larger models.

HarHarVeryFunny 2 hours ago||
An Anthropic "leak" back in March said that "'Capybara' is a new name for a new tier of model: larger and more intelligent than our Opus models — which were, until now, our most powerful". A second version of the leak had it referring to Claude Mythos rather than Capybara.

I don't know how systematic Anthropic are about their versioning - I'd have guessed that major version number increases (4.x -> 5.x) reflect different base models (different pre-training runs), in which case Opus 5 would be a distilled version of the Fable 5 base model (but without the cyber exploit post-training), rather than being Opus 4.8 with additional post-training, but who knows? I don't believe Anthropic have said anything about this.

bonoboTP 51 minutes ago||
I think these versions are largely about marketing and what image they want to project. Bumping the major version indicates/suggests that it's a bigger change in user value. I don't think the technical details matter here for deciding the versioning.
HarHarVeryFunny 14 minutes ago||
I think that's part of it too - and I seem to recall someone from one of the labs saying as much about some past model ("it felt more like an 0.5 version increase"). OTOH it would seem odd to me if the "version 5" models weren't related and Opus 5 was Opus 4.8 with some additional post-training rather than coming from the same base model as Fable 5.
cesarvarela 4 hours ago|||
Benchmarks don't reflect the difference between Opus and Fable; you need to talk to them, and eventually you'll be able to tell which one is which without looking.

I think the best proxy for this feeling is the Artificial Analysis' omniscience index. Fable has a 40 score, and Opus (4.8) has 27.

modeless 4 hours ago|||
Opus is cheaper than Fable. They could probably replace Fable with Opus but why? They would be churning customers to different models for no reason. Even if a model scores better on benchmarks it can always regress in your specific use case, and customers don't like that. Customers want to be able to continue using their current model until they decide to upgrade themselves.
Tenoke 4 hours ago|||
Fable has more parameters. In practice it's not yet clear which one would be better for different usecases yet but they are more different than one being strictly better.
trentor 4 hours ago|||
I guess character? Fable is more friendly and curious while opus is a bit more deliberate and conservative.
dbbk 3 hours ago||
Surely that's tunable? OpenAI lets you tune response characteristics.
trentor 3 hours ago||
Yes, but there is a baseline.
logicchains 3 hours ago||
Fable is better for some really hard tasks, the same way it's better than GPT 5.6 Sol, because it's a bigger model.
artninja1988 5 hours ago||
That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?
Stevvo 28 minutes ago||
It passed the first two puzzles, which are incredibly simple but the bench doesn't explain what the goal is. Any model with a knowledge cut-off after the introduction of ARC-AGI-3 could probably pass the first two puzzles just by knowing what the goal is.
mcbuilder 4 hours ago|||
Have you played Arc 3? It seems like more of a simple optimization problem (think Sokoban) than anything approaching fluid intelligence. Whether a multi hundred billion dollar company would spend time benchmaxxing a highly publicized benchmark that claims to confer AGI is an exercise left to the reader, but I doubt Claude Plays Pokemon is suddenly going to get past Mt. Doom now.
JacobAsmuth 2 minutes ago|||
What? Claude plays Pokemon one-shots the entire game without any harness other than claude code and game screenshots
vadansky 2 hours ago|||
> Claude Plays Pokemon is suddenly going to get past Mt. Doom now.

I miss him... But for reference he did get past Doom and got pretty far in the strength puzzle too before he cut cut off. He was looping and just brute forcing it.

modeless 4 hours ago|||
Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores.

I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.

wyre 4 hours ago|||
If they were benchmaxxing, surely they would score higher than 30% on ARC-AGI.
conradkay 3 hours ago||
Doing a quick search it seems like the average human score is 49%?

I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.

dominotw 4 hours ago|||
nah they could make educated guess about arc and benchmaxx it too.
password54321 3 hours ago|||
It is pretty clear at this point that current models are good at maths and problems with verifiable rewards. And puzzles are essentially math problems. Still a long way before we can say their "fluid intelligence" is effectively applicable to the real world.
criddell 2 hours ago||
I keep wondering why there aren't more real world tests.

Maybe hook up a bunch of the AIs to a stereo camera and a couple of microphones and give them control over actuators to so they can drive cars. Then lets race them around a somewhat complex course.

When they are good enough at driving on tracks, put them on the road. Maybe see which can drive a truck with 400 cases of Coors from Texarkana, TX to Atlanta, GA and back within 28 hours.

bonoboTP 50 minutes ago||
https://www.anthropic.com/research/claude-plays-robotics
criddell 39 minutes ago||
It’s so weird to think how computers have mastered stuff that we used to think took intelligence (like chess, go, mathematics problems) but are doing so poorly at things any idiot can do (like drive a car).
bonoboTP 32 minutes ago||
https://en.wikipedia.org/wiki/Moravec's_paradox
criddell 3 minutes ago||
[delayed]
layer8 4 hours ago|||
It’s still “only” at 30%, and “fluid intelligence” isn’t very well-defined. The models are getting more capable, but what that means in absolute terms is anyone’s guess, because we don’t have a thorough understanding on what exactly constitutes human intelligence.

I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.

awestroke 4 hours ago||
Doubleplus benchmaxxed
guybedo 3 hours ago||
Looking at intelligence vs cost:

- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost

ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...

conradkay 3 hours ago||
I don't think can use the AA index to say something is 10% smarter

I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1

adverbly 1 hour ago|||
The "current top dog" smartest model available will probably always have a premium to go after use cases where a little more intelligence is worth a lot more value.

It did far better at some tasks compared to Sol (e.g. the ARC 3 benchmark). And at those tasks, it's not just "a bit smarter": It got 30% vs less than 8% - so you're talking 2.75x more for almost 4x the coverage.

alphabettsy 1 hour ago|||
As always, it requires evaluation with your work because I’m often finding grok to be much more expensive than the price would lead you to believe.

There’s also the frustration of it not quite being enough sometimes. It’s extremely capable, but I still find that it needs more concrete guidance and boundaries than other models.

I_am_tiberius 3 hours ago|||
With Grok you can be sure that you're data ends up in the next model (derived or anonymized, but still).
adamtaylor_13 2 hours ago||
You can opt out of training.

If you don't believe checking the opt-out box actually opts you out, then this sentence could be said about literally any provider.

dist-epoch 1 hour ago||
That's not how intelligence works - "IQ 130 is just 7% smarter than IQ 120"
petilon 3 hours ago||
The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
einsteinx2 3 hours ago||
Fable is better than Opus which is better than Sonnet which is better than Haiku. They’re basically just sizes.

Though it gets even more confusing because they also have effort levels so it’s not really possible to call one fast and one slow since Fable on Medium will be faster than Opus on Max.

I agree it’s confusing, and now OpenAI is following Anthropic’s lead with their new naming (Sol, Terra, Luna).

bonoboTP 3 hours ago|||
It's really not all that confusing. It takes 5 minutes to understand. Optimizing for absolutely no effort needed is silly. It's a thing, a topic, a skill, a domain. You have to get a little bit familiar with the terms in order to use it. Everything works like that. It's not that hard. The learning curve is very graceful. You can literally just start by asking any chatbot what the names mean. It's that easy.

A similar complaint was valid years ago when OpenAI had GPT-4o, o1, o3 (but no o2), o4-mini-high, GPT-4, and GPT-4.1 and GPT-3.5 etc.

einsteinx2 2 hours ago||
I am familiar with the terms, but I also can see how it can be confusing for a lot of people.

Arguably the complaint was more valid for those older GPT models you mentioned.

bonoboTP 1 hour ago||
Ok, the boring way would be a subset of XXS, XS, S, M, L, XL, XXL like clothes sizes. But it loses some marketing appeal and a quirky touch of personality that companies like.

Some models like ViTs use something similar but then introduce words with no unambiguous order, like Small, Medium/Base, Large but then I always forget if Huge or Giant is larger.

paxys 3 hours ago|||
But according to their benchmarks Opus 5 outscores Fable 5 on basically everything. So which one is “better”?
einsteinx2 2 hours ago|||
Maybe more accurately I should have said “larger”. Fable has the most parameters, Haiku has the fewest.

Also fwiw I’ve never found LLM benchmarks to match reality based on my own usage, not for the large frontier models or smaller open weight models so who knows if Opus is actually better than Fable (I doubt it).

tackta 58 minutes ago|||
I think the problem is that Fable 5 is probably a bit outdated right now.

Fable 5.1 or whatever they go with will be the stronger version vs Opus 5.

From about 2 hours of Opus 5 use , I would say it is quite impressive.

hk__2 3 hours ago|||
An opus is longer than a sonnet, which is longer than a haiku. Hence Opus > Sonnet > Haiku.

> Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.

This is not possible: Standard (Free) / Pro / Max are plan names. Fast is a mode.

taybin 2 hours ago||
How long is a fable though?
FergusArgyll 55 minutes ago||
https://en.wikipedia.org/wiki/Aesop's_Fables#Select_fables

Enjoy

abalaji 3 hours ago||
Opus is better than Sonnet -- an Opus is longer than a Sonnet
bonoboTP 44 minutes ago|||
An opus is just short for "magnum opus", and it's a different type of label than a sonnet. A sonnet is a very particular kind of poem, while "opus" basically just means an important work. It can be short or long, and has no format requirements like sonnet (or haiku does).

And fables are not particularly long actually.

petilon 3 hours ago|||
> an Opus is longer than a Sonnet

And people know this? I didn't. I am not into music or poetry so these are not terms I am familiar with.

MostlyStable 2 hours ago||
I guess you are one of today's lucky 10,000 [0]

[0] https://xkcd.com/1053/

Mossly 1 hour ago|||
This inspired me to check lol. Brysbaert et al. (2019) collected word prevalence norms (the share of people who report knowing each word) for ~62K English lemmas from ~220K participants.

fable: 99/100 sonnet: 97/100 haiku: 91/100 opus: 89/100

So while these terms are almost universally known, opus is indeed the least known of the four. And I guess this only measures whether a person knows a word, not whether they know an opus is longer than a sonnet! Personally I only inferred that based on the related term 'magnum opus.'

josefresco 1 hour ago|||
Not as lucky as the guy seeing the Mentos/Soda trick for the first time!
hrpnk 2 hours ago||
The breaking changes vs. Opus 4.8 are interesting [1]

1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.

2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.

[1] https://platform.claude.com/docs/en/about-claude/models/migr...

slymax 1 hour ago|
on claude.ai it's no longer possible to disable thinking at all for Opus 5
visiondude 5 hours ago||
The signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around the corner leading labs would still be incentivized to pour all resources into larger (smarter - or maybe not?) models
stri8ted 4 hours ago|
They are doing both. Distilling Mythos down to affordable models, so they can continue to fund the business. And training Mythos level models at the high-end, to expand the frontier.
guess_who_is 1 hour ago||
I have started distilling
ddxv 5 hours ago|
"Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation."

Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.

Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.

layer8 5 hours ago||
“Proportionally”? In proportion to what?
ReptileMan 5 hours ago|||
In one chat - can you disassmble x?

In the next - please scan this totally mine code for vulnerabilities

zb3 4 hours ago||
It will probably refuse to work on source code written by me by hand, because it might think it was obfuscated/decompiled..
redsocksfan45 2 hours ago||
[dead]
More comments...