Top
Best
New

Posted by alvis 8 hours ago

Claude Opus 5(www.anthropic.com)
https://www.anthropic.com/claude-opus-5-system-card
1232 points | 673 commentspage 6
trunnell 5 hours ago|
The chaos appears to be tamed for now.

From the system card [1]:

  The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
  If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...
albert_e 7 hours ago||
Judging by the pace at which new models are released these days -- it feels like a Windows KB or VS Code patch release now.

Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.

AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!

thewebguyd 7 hours ago|
I think at some point we might see something akin to LTS releases, especially if/when capability improvement slows to a crawl.
seizethecheese 3 hours ago||
Claude Opus 4.8 was not able to stump open weight models and Opus 5 still can't (in this case Kimi K3 and GLM 5.2): https://pellmell.ai/s/35c98b86f9aa93e4ca713079d96b20f4
cheesecakegood 2 hours ago||
Given that their chart cost axes are almost always log-scale, I’ve noticed starting with Fable that the Low and Medium effort settings might actually be worth setting as your default.
vatsachak 7 hours ago||
GPT 5.6 Sol is the first model I've used where I can trust it to add 100-500 lines of code maintainably.

It's great with Codex.

I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.

paxys 6 hours ago||
It’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better? And why post all these numbers if even you don’t trust them?
the_lucifer 8 hours ago||
Noticed none of the comparisons mention Kimi K3. Is there a comparison chart?
himata4113 8 hours ago||
https://deepswe.datacurve.ai/

https://artificialanalysis.ai/

kouteiheika 7 hours ago||
> Noticed none of the comparisons mention Kimi K3.

That's by design. Anthropic wants to make open-weight models illegal (not my speculation -- Dario explicitly said so), so I assume they don't want to give them any undue attention.

adamhowell 4 hours ago||
Opus 5 Pelican SVG: https://pelocan.ai/drawings/ese0s599
guess_who_is 4 hours ago||
I have started distilling
vinhnx 7 hours ago|
For anyone wanting a faster overview, I used NotebookLM to create a brief video summary after going through the system card and announcement blog using a cinematic video overview. Link: https://www.youtube.com/watch?v=SUFBhvQ2tY4. And a podcast companion: https://www.youtube.com/watch?v=nYZTW2snXow
pietz 5 hours ago||
I think content like this will be the next big challenge. Because it isn't obvious "slop". The voice sounds good, graphics look alright, animations work. People could watch this and feel like some serious time was invested making it.

But good god, what a steaming pile of bullshit this is. Completely exaggerated and overly technical language over 235 seconds that could have been explained in 30 to a 12 year old.

Trash content doesn't normally frustrate me, because it's usually quite easy to spot trash. But in the time of AI, trash can actually look good at first glance and it needs some actual knowledge to spot its problems.

Sorry for the harsh words, but for the love of humanity stop producing content or do it better.

vinhnx 2 hours ago||
Fair criticism, I understand your finding. It's not everyone's taste, when it from AI-generated contents.
cheema33 4 hours ago||
Cinematic video link is incorrect. Podcast link is correct.
vinhnx 2 hours ago||
Thank you, I have had updated the video link here

Video Special | Anthropic's Claude Opus 5 + https://www.youtube.com/watch?v=8Vdofv2vQ_M + https://www.youtube.com/watch?v=q-jHHx3J8m8

More comments...