Top
Best
New

Posted by alvis 10 hours ago

Claude Opus 5(www.anthropic.com)
https://www.anthropic.com/claude-opus-5-system-card
1325 points | 717 commentspage 8
dehugger 10 hours ago|
Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?
andrewl-hn 9 hours ago||
I suspect they make a big model first. In this case it's Fable. Then they run the shrinker steps to make Sonnet and Opus. Sonnet is smaller, takes less time to make, so it got released first. Opus needed few more weeks to cook.

With this iteration they had a delay because when the Mythos was ready they had some sort of "Oh shit" moment and spent half a year adding safety guards to it. Then slowly rolled it out, but got another delay due to a government block. So, maybe the work on making Opus and Sonnet only started after they got a green light from the administration.

Presumably, now that they learned how to do this safety-wrapping the next iteration of Mythos / Fable / Opus / Sonnet is going to show up faster.

Something like that.

Wowfunhappy 9 hours ago||
But I wonder how they were able to release Sonnet 5 during the period when even people inside Anthropic were legally barred from using Mythos/Fable?
anon373839 5 hours ago|||
There is word on the street that they allowed employees to work with a slightly stronger internal version of Mythos (5.1, if you will) that, in their parsing of the order, wasn’t restricted. If true, in practical terms, they ignored the order.
riknos314 9 hours ago|||
Iirc the ban only applied to non-Americans. While anthropic found collecting citizenship information on all customers too burdensome, it's a much smaller lift to collect such info for your own employees.

So I'm assuming at least a subset of employees could continue using the models during that time.

Wowfunhappy 8 hours ago||
Although the ban was only for non-Americans, Anthropic said that they'd also restricted access to their own employees internally, because they had no other realistic way to apply the government's orders. I guess it's possible they were lying, but seems unlikely.
tedsanders 8 hours ago|||
No, they're different models. Knowledge cutoff has been updated.
oh_no 9 hours ago||
based on pricing I think it's safe to say they're different. why would they charge half price when fable has been very popular?
somenameforme 8 hours ago||
They just got a huge amount of customer price info over the past few days after they went token only for Fable. I suspect the conversion rate was extremely low, with consumers far less sticky than they might have hoped. In my case I was planning on swapping, probably to a Chinese model, when my sub expired this month, but the release of Opus 5 is probably enough to keep me paying rent until the next open model/closed model face-off in a couple of months.
6thbit 9 hours ago||
"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them."

"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".

This is probably great news, but then again, where does this leave Fable as a choice?

mulhoon 8 hours ago||
As a coder, I’ve had no desire to use Fable. In fact I switched from Opus models to sonnet 5 and haven’t noticed any drop in quality on large repos. It seems the gap at the top is very small and not hugely noticeable for backed/frontend. Has anyone else had this experience?
furyofantares 7 hours ago||
If I'm using medium or low reasoning, I use Sonnet 5. If high or above, I use Opus 4.8. (Before 5, I was never using Sonnet. This is a Sonnet 5 vs Opus 4.8 comparison.)

Sonnet 5 and Opus 4.8 seem about the same to me - the reason I switch between the two is I'd read that it's cheaper to use Sonnet 5 on those reasoning levels, and cheaper to use Opus 4.8 above them. This is due to them using different token quantities.

stsch 7 hours ago||
I use Opus for specs and planning, Sonnet for code generation.
itissid 7 hours ago||
I found opus 4.8 too agreeable and too wordy(as opposed to codex) and too agreeable. If you are reading documents generating by it was too much. TBH. Fable did a bit better on this. Anyone seen a marked difference with opus 5 on this?
ianberdin 7 hours ago|
Opus yes, it likes to explain steps and reread files.

Fable is not better, it says zero information between steps and then output a summary. A perfect “send - done”.

firemelt 2 hours ago||
so what is the default effort for this model?
markasoftware 10 hours ago||
Soo most of the benchmarks are better than fable... Is this naming scheme just to avoid getting banned again?
geooff_ 10 hours ago||
FYI: `/model claude-opus-5` works to use it even through `/model` still tries to serve 4.8
dpe82 9 hours ago|
`claude update`
6thbit 9 hours ago||
Anyone has an insight into how much money labs are putting into benchmarks?

Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..

jofzar 33 minutes ago||
It's $0/close to 0, they aren't at 100% demand so any leftover compute is "not spent".

The cost they are quoting is API cost, so it's already inflated on that.

stri8ted 9 hours ago||
20k is small potatoes for the marketing impact.
vinhnx 8 hours ago|
The benchmark appears to have a mistake, as Opus 5 and Fable 5 score 53.4% and 53.5%, respectively, for the Agentic Coding row (FrontierCode v1.1). But Opus 5 is the highlight.
bouke 8 hours ago|
How hard can it be to be to correctly annotate the table? DeepSWE doesn’t have a highlight either; Fable slightly better than Opus (69.7% vs 68.8%).
More comments...