Top
Best
New

Posted by alvis 11 hours ago

Claude Opus 5(www.anthropic.com)
https://www.anthropic.com/claude-opus-5-system-card
1361 points | 737 commentspage 9
alvis 11 hours ago|
What really impress me is opus 5 is better in alignment than fable 5!
briandoll 11 hours ago||
Very interesting to see such a focus on cost for performance here
m_w_ 11 hours ago||
Very impressive headline benchmark numbers. I expected a step change, but not past Fable. That said - it all depends on whether the classifiers make the model unusable...
tyre 10 hours ago||
I'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.
somenameforme 9 hours ago|
In what ways have you found it better than just typical code based UI iteration? Considered checking it out but never really got around to it as I'm generally okay with Claude's UI work so far.
pmg1991 11 hours ago||
Same cost as 4.8 but better that 4.8. Happy to get more efficient model. But is there any reason all companies are releasing models back to back after GLM 5.2.
aleenz1102 10 hours ago|
"Bro, AI model releases have officially overtaken iPhone releases. At this rate, we’ll be getting 'Claude 9.0 Extra Crunch' by next Tuesday."
twothreeone 11 hours ago||
It starts at page 148.
boc 10 hours ago||
Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:

"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.

What I can tell you is what I actually observe:"

I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.

One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!

jannyfer 10 hours ago|
So wordy.
firemelt 3 hours ago||
so what is the default effort for this model?
bovermyer 10 hours ago||
This stood out to me as a little concerning:

> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.

orangecat 10 hours ago|
That seems to be for the "AA-Omniscience" test where you get +1 for a correct answer, -1 for a wrong answer, and 0 for "I don't know". If a model is more than 50% confident in its answer, it should go ahead and submit it even though it will sometimes be wrong.

I'd be curious to see a version of the test where models are asked to give a probability that their answers are correct so we can see how calibrated they are.

skinfaxi 11 hours ago|
Related https://news.ycombinator.com/item?id=49038393
More comments...