Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.
Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.
---
I provided crapton of context for that one resume line. All the work I did, documentations for my justifications, etc.
I initially messed up and came out ot $5, rest of resume used around $4 per line (I used a fresh new session on purpose).
---
As a clarification, $2.2 average for OPUS 5.5 was the same process in a new session, same context, same prompts.
Also adding verification for that Fable 5.1 output in the same sesssion.
Nice. I was starting to think that Haiku got abandoned.
I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.
All of this starts to feel more like a drug dealer selling their newest stuff.
In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.
And on the way I always have to check my tooling and need to adjust things to get max results.
Now, Anthropic might stall on releasing Fable 5.5, due to the "pacing the frontier" threat-to-humankind management business. If so, Fable 5.1 would remain a niche model for the next bit.
Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.
Benchmarks often don't survive contact with reality.
Yes, in the sense that it reproduced results in the paper or known solutions obtained by other methods. In fact, Opus is very good at checking it's own work in my experience.
Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Opus 5.5's output: https://html.non.io/annui-opus/
Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.
For comparison with other drops this week + current #1:
Astra: https://html.non.io/annui/
MiMo: https://html.non.io/annui-mimo/
Grok 4.7: https://html.non.io/Annui-grok/
Worth noting though that GLM 5.3 isn't multi-modal, so it doesn't have a vision layer. It is quite clever and hacks around it pretty effectively however. I'm running a deepseek 4 build now and will reply shortly with that.
The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.
Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.
For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.
The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.
One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
Can you give more details here? This sounds intriguing.
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
https://artificialanalysis.ai/models/releases/claude-opus-5-...
Opus 5.5 Medium = $1.34
GPT-6-Astra High = $1.76
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference): Opus 5.5 High = $1.82
GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
That's an incredibly bold assumption.
This is news to me. Excited to try it out! Thanks.
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
OSWOrld - XLANG Lab @ the University of Hong Kong
Chartography - Surge AI
"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"
and
"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."
and
"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."
I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.
If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.
Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.