Top
Best
New

Posted by alvis 7 hours ago

Claude Opus 5(www.anthropic.com)
https://www.anthropic.com/claude-opus-5-system-card
1173 points | 638 commentspage 5
lucamark 6 hours ago|
But why GPT 5.6 Sol is so behind on the benchmarks? In real-world projects, it is the best frontier model to me in terms of accuracy, speed and consistency. It can just be compared to Fable 5, but I prefer GPT 5.6 Sol because of inference speed.

I've never trusted on model cards though. I'm sorry.

dbbk 5 hours ago|
Another benefit is that fast mode can be used on subscription, but Anthropic's won't
lucamark 5 hours ago||
Exactly! And they should also release the new inference engine in this month. Anyway, I am curious to try Opus 5, considering that previous versions (e.g., 4.8) were disappointing
modeless 6 hours ago||
Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
oh_no 6 hours ago||
seeing a jump this big is not a great sign for the continuing value of a benchmark
modeless 6 hours ago||
It will continue to be valuable as a cost and speed benchmark long after it is saturated at the high end. And they are already working on ARC-AGI 4 and thinking about going even farther.
dominotw 6 hours ago||
> I continue to believe ARC-AGI measures something different

why is that? its now being benchmaxxed too

williamstein 6 hours ago||
> This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.

Annoyingly, this is a concrete argument that open source software may be easier to attack.

seizethecheese 2 hours ago||
Claude Opus 4.8 was not able to stump open weight models and Opus 5 still can't (in this case Kimi K3 and GLM 5.2): https://pellmell.ai/s/35c98b86f9aa93e4ca713079d96b20f4
trunnell 4 hours ago||
The chaos appears to be tamed for now.

From the system card [1]:

  The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
  If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...
irthomasthomas 6 hours ago||
Changelog - fixed issue where model acts like qwen when prompted in chinese
bottlepalm 6 hours ago||
Page 151 of the linked system card - did Opus 5 get nerfed to prevent it being better than Fable? The graph makes no sense. Huge decline in coding performance at effort levels higher than medium.
luciana1u 4 hours ago||
473 comments in 3 hours. people are speedrunning having opinions about it
albert_e 6 hours ago||
Judging by the pace at which new models are released these days -- it feels like a Windows KB or VS Code patch release now.

Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.

AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!

thewebguyd 6 hours ago|
I think at some point we might see something akin to LTS releases, especially if/when capability improvement slows to a crawl.
adamhowell 3 hours ago|
Opus 5 Pelican SVG: https://pelocan.ai/drawings/ese0s599
More comments...