https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Here's how the thinking effort levels compare:
low
27 input, 1,623 output, thinking_tokens: 0
1.6284
Duration: 10138ms (10s)
medium
27 input, 1,796 output, thinking_tokens: 0
1.7914 cents
Duration: 11266ms (11s)
high
27 input, 2,334 output, thinking_tokens: 745
2.3394 cents
Duration: 17376ms (17s)
xhigh
27 input, 5,730 output, thinking_tokens: 2535
5.7354 cents
Duration: 41882ms (41s)
max (failed to return response)
27 input, 128,000 output, thinking_tokens: 128000
$1.28
Duration: 940617ms (15m 40s)
Low and medium both used 0 thinking tokens.2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: https://jonclegg.github.io/pacman-bakeoff/
About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like "the blaster weapon is way too powerful, divide it's hit points by 10" or "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS")
Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…
Any model release it’s the top comment, I do not understand why.
In the GPT-6 comment I included full visual comparison grids: https://news.ycombinator.com/item?id=49805509#49806126
For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: https://news.ycombinator.com/item?id=49639090#49645591
30% chance of responding with something about Enshittification and how it can't fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it's running on FreeBSD behind PiHole).
30% chance of complaining that it's being subsidized and that "prices are going to go up bro."
30% chance of some unrelated rant on ID checks for age verification.
10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.
PS: the next human that brings up pelicans on bicycles should try to draw them.
I’m not sure whether that’s a feature or a bug at this point though.
I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.
That would make it about so, I assume?
AA Output Reason Cost
Kimi K3 Max 44 48k 32k $2.00
Half tokens 44 32k 16k ?
Opus Med 51 26k 12k $1.34
Opus High 54 36k 18k $1.82
Opus Max 58 119k 84k $5.98
Sonnet Med 41 ? ? $0.59
Sonnet High 47 ? ? $1.08
Sonnet Max 56 193k 142k $7.60
Medium is Anthropic's default.Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.
Gemini 3.8 Flash is 65,536 https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flas...
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I mostly use Fable though, Opus only via sub-agents.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
There was also a period where specifically OpenAI models would always have to comment something in code review and the builders were agreeable up to listening to each nitpick. If you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about. Tried it this week with Astra reviewer and it's about 0-2 review loops (never had a LLM accept a change without nitpicking first try before Astra).
There was also a period where you'd have to give quite specific instructions for agents to keep iterating, but now agent are pretty proactive and try to finish tasks you give them unsurprisingly most of the time.
So, while there's a shortcoming of LLM+harness and engineers observe more tokens improve things even logarithmicly, you'll see more tokens seemingly abused by engineers.
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
[1]: https://code.claude.com/docs/en/channels [2]: https://code.claude.com/docs/en/channels-reference
But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
Could you explain why it would be a goal to understand the system less, rather than more?
It seems harder to know if you have good tests while lowering your expertise in the system.
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.
Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".
I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.
OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.
Totally different uses.
In short, seems to describe vibe-coding to me? What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.
There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):
> I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.
> The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.
> The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.
What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)
I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.
I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
[0]: https://news.ycombinator.com/item?id=49808422
[1]: https://sockpuppet.org/blog/2026/09/25/what-even-is-an-os-no...
For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.
Why hire a plumber when you can just watch some youtube videos and do it yourself?
Why pay someone else for their software when you can just make your own?
Because the hard part of making software wasn't *just* writing the code. It was about understanding the problem well enough to understand what the solution should look like.
I feel like as software engineers we should be pretty familiar with what it's like talking to your average user, they will sometimes understand the root cause of what's making their task difficult (although often will get focused on some annoying but ultimately trivial symptom) and have very disasterously bad ideas on how to solve it.
What we've given them with generative AI is a machine they can put their sometimes ok, sometimes questionable understanding of the problem and their dreadful solutions and it will happily churn away building it regardless of how pointless and silly it is.
A future where every user can tell the AI "We keep getting the sales tax wrong, remove charging sales tax from the checkout flow" isn't one I'm terrifically worried about.
In the same way that having access to information about plumbing didn't suddenly make everyone plumbers, having access to a machine that will implement every idea you have regardless of quality doesn't suddenly make everyone a software engineer.
There are a lot of horrible potential scenarios that are really scary to contemplate. There are also a lot of really delightful ones where AI does the drudge work, invents a million incredible medicines, and frees us up to hang out and make art all day. And there are even more scenarios somewhere in the middle where AI changes a lot of stuff but we all still more or less end up going to work and doing jobs.
I've basically had a background thread in my skull running at high priority for the past two years trying to predict which of those scenarios I think are most likely so that I can plan for them. It is utterly exhausting spending that many mental resources on a question like that.
It finally clicked for me a couple of weeks ago that no one is going to be able to accurately predict all the thousands of ways AI will affect the world. Certainly not me. We are living in unprecedented times. No one has a map for the future.
So I am trying to loosen my hold on the future some and focus more on the present. I have a great job and a great family now. I have most of my health. I'll try to live my life right now to the fullest and in accordance with my values. The future is going to have to be future me's problem. That's OK.
The proof of the pudding.
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
The speed of your manual reviews become the limiting factor, which you should be doing at some level to maintain sanity, even if there are enough ideas to be worked on to maintain a review queue.
Not quite sure where this fits well. Maybe small one one off requests like using Claude desktop/web?
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
> And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding.
Yes. Our career is over, as is our economy. Soooo... FYI :(> the economy is over
Hackernews' neuroticism remains undefeated
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
Damn, HN commenters starting to talk in claudisms now
this is claude writing...
corporate needs you to find the difference
Anthropic made it that way, and I'd say the lower score is accurate.
> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).
> https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...
I cannot find a Sonnet 5.5 system card.
AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k
Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.
I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
Daybreak Blue is the not the same thing as Daybreak Red, which has a more significant hurdle. I don't know anyone who has gotten access to Red.
And my job won’t even pay for Claude now because it’s so ruinously expensive.
Mimo 2.6 Pro: 0.04/0.4/0.87
Sonnet 5.5: 0.2/2/10
Opus 5.5: Sonnet prices times 2
What I dont understand is their cache writes ($2.5). Why is that not covered by input cost?
The pricing model confuses me though (I presume by design, Hanlon be damned).
OpenAI and Anthropic have both transitioned into product companies. ChatGPT (the app) and Claude are both one-click installs that just work. People and businesses with pay for this.
People will also pay for the best (or the perception of being the best). Since it's hard to tell what "intelligence" really means model to model, there's a sense of safety in giving a task to the "best".
Microsoft has been releasing dog shit insanely overpriced software with decent alternatives for decades and is still used in every single company I work for or with.
Your take is the "current year is the year of the linux desktop" meme of "ai"
I don't think anything comes close to Microsoft's offerings. Macs suck. Ditto Linux.
People use them, if for no other reason, because they are cheap, or are part of the Chromebook generation and have gotten used to it
Of their suite, Presentation and Sheets are the only ones people really have gripes about, Sheets by power users because it isn't Excel and it can never be, and Presentations because it's the ugly duckling of the suite
Even 5.2 is doing really well in comparison here: https://labs.scale.com/leaderboard/sweatlas-refactoring
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
i think they see what openai charges for luna and just don't want to try and compete
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
This is bollocks. Their safeguards are shit.
https://support.claude.com/en/articles/14604842-real-time-cy...
> This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We'll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models
You obviously should not expect the CVP to cover this model either.
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
This is the way.
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
In general I am sympathetic to the argument that a chat interface can't really distinguish between white hat and black hat pen testing, but it seems absurd to have a verification program if it doesn't skip most of those checks.
The silicon valley ethos is "ban early and often, and invest nothing in appeals systems", so any gate before that helps!
AI providers still haven't realized how much cash they could rake in if they provided fully unrestricted models.