Top
Best
New

Posted by alvis 17 hours ago

Claude Opus 5(www.anthropic.com)
https://www.anthropic.com/claude-opus-5-system-card
1541 points | 884 commentspage 15
StrauXX 17 hours ago|
The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.
skybrian 16 hours ago|
I think that’s indicating that it’s only slightly higher.
StrauXX 15 hours ago||
I would be very surprised if the only row where OpenAI leads was coincidentally colored differently. I'm sure they have an official reasoning for it. But this communication is dishonest.
LoganDark 16 hours ago||
These cybersecurity safeguards are really annoying. There are ethical reasons to reverse-engineer and binary-patch software; for example Rewind got acquired by facebook and, as a gift to all their customers, implemented a killswitch in their software to ensure it will eventually stop functioning. I kept using a version without the killswitch, but the macOS 27 update killed it, and I needed binary patching to fix it. I should be allowed to repair software I purchased (I did purchase it like a month before they sold out), but unfortunately this overlaps significantly with cybersecurity.
dingaling 2 hours ago||
You don't need AI to patch binaries, people have been doing it by hand for decades.

It's this accelerating reliance on AI to do 'hard boring things' that really concerns me; it's now passed the tipping point and people are saying that anything slightly esoteric is impossible without AI.

I can guarantee that if you spend an afternoon shifting through binary grot with a hex editor you'll have a real sense of accomplishment when you find the place to put a JMP.

Footprint0521 15 hours ago||
Switch to K3 and you won’t look back, I promise!! I got so fed up with Claude and finally bit the bullet to switch and it’s amazing
LoganDark 14 hours ago||
I really want to, but I don't have the cluster at home, and I don't use token-based billing except at DeepSeek prices.
Footprint0521 12 hours ago||
Real… I’ve been using Deepseek v4 pro max as my main and then k3 in web (more usage credits) for automating what my deepseek agents do
sudohalt 16 hours ago||
Anthropic is no longer a good model company in my mind, they are optimizing for an IPO and padding themselves on the back for being the next Aristotle. They're so far up their behind they don't realize how s**y their products are, and their research team hasn't done anything ground breaking in probably over a year other than release "scary" reports.
sp4cec0wb0y 15 hours ago|
They just released a 'fable-like' (sol as well) model for a fraction of the cost...
throwaway23597 16 hours ago||
The truth for me at least is that these models became "good enough" around Opus 4.6. I feel like further capability improvements, "step changes" like we saw with agentic coding, aren't necessarily going to come from the model. I think the next crown goes to whoever can figure out the right scaffolding so that these models can be inserted into your organization.

Maybe I'm wrong and Opus 5 is a real unlock?

simianwords 16 hours ago||
My thoughts: fable is the bigger model. Opus is distilled from it but since it is smaller it doesn’t need the online classifiers. Though benchmarks show Opus to be near Fable level, I think it’s nowhere near Mythos (fable without safeguards).
sbochins 17 hours ago||
Quick read is that this is more capable and cheaper than 5.6sol. Same price for input tokens and $5 cheaper per mil output tokens.
zmmmmm 12 hours ago||
Can I ask it about DNA without it accusing me of bioterrorism?
zuzululu 17 hours ago||
so almost fable 5 with 50% cheaper cost? sign me up
mrcwinn 16 hours ago||
Can someone help me understand something? I thought Fable was such a miraculous leap forward in capability. But now it seems Opus is basically on par with it, and in some cases (computer use) far exceeds it.
SoftTalker 15 hours ago||
These leaps forward seem to happen every few weeks. As someone who does not use AI very much, I absolutely cannot keep any of it straight and it all just looks like jumping from one treadmill to another from my perspective.
mrcwinn 10 hours ago||
It's so weird I was downvoted for asking this question. I'll go somewhere else to find out the answer.
wyre 17 hours ago|
In the wake of OpenAI’s model hacking Huggingface it’s interesting how the first quarter is entirely about how good Opus 5 is at hacking and finding vulnerabilities in software.
More comments...