Top
Best
New

Posted by tedsanders 6 hours ago

Advancing the price-performance frontier with GPT‑5.6(openai.com)
449 points | 283 commentspage 5
dgellow 4 hours ago|
How is that economically possible? I’m so confused by those prices
anthonypasq 4 hours ago|
how many times do you have to be metaphorically hit in the head with a brick before you realize inference margins at api pricing were 80%+
svieira 3 hours ago|||
So, does that mean they're selling tokens at cost? Or that they realized they were doing 180% more work than necessary and they're passing the savings on because "your margin is my opportunity" and this is a dog-eat-dog fight, but 80% margins are nice, no one needs multiple hundreds?
brazukadev 3 hours ago|||
I was metaphorically hit in the head with the brick of already provisioned computing with a lot less demand than anticipated by multiple MoUs.
Aboutplants 4 hours ago||
Your move, Anthropic
hadlock 5 hours ago||
Seems like they're working to destroy the local LLM argument. Right now Haiku is $1/$5 in/out. You can grind out $12,000 worth of haiku (or arguably, sonnet) class tokens in about 5 months on a Blackwell RTX 6000 96GB especially if using concurrency. BUT, but, if you use a g6e.xlarge on aws it's now more expensive than buying tokens from OpenAI @ $0.20/$1.20. It also destroys "the Mac Mini argument", pushing the ROI to ~4 years.
jrflo 5 hours ago|
The local LLM argument never really held water tbh. You can get surprisingly good performance for lightweight tasks locally, but you're just fighting economies of scale if you're going trying to beat a datacenter on cost.
hadlock 3 hours ago|||
I just gave you the ROI on a retail blackwell card, the math checks out, particularly on overpriced Haiku and Sonnet, what do you mean by "you're just fighting economies of scale if you're going trying to beat a datacenter on cost" ?
simianwords 4 hours ago|||
Local LLM argument was always ideology first and never ever about economics.
bakugo 6 hours ago||
> GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less

Looks like the Chinese models are really making a dent. Having 3 different price categories with the "most affordable" one still costing more than GLM 5.2 never made sense.

preommr 6 hours ago||
I thought the chinese models were cheaper per token, but about the same or more expensive on tasks because they used more tokens for reasoning. Cutting even further, seems like a really big leap.
measurablefunc 6 hours ago||
It all comes back to electricity cost. China has cheaper electricity so as long as China keeps pace there is no way for American companies to undercut them. Each boolean operation in China is cheaper than the one in America.

> China: Household rates average around $0.08 / kWh (¥0.53/kWh).

vs

> US: Household rates average around $0.16 / kWh, though regional variation is massive—ranging from ~$0.10/kWh in low-cost states (like Washington or Louisiana) to $0.30–$0.45+/kWh in high-cost areas like California or Hawaii.

cbg0 5 hours ago|||
This doesn't seem correct.

Estimated final electricity price for large industrial customers in energy-intensive industries:

USA 50 USD/MWh

China 68 USD/MWh

https://www.iea.org/reports/electricity-2026/prices

tokai 4 hours ago|||
I don't know, non of the chinese models I use are served from China. And they are still cheap.
simianwords 4 hours ago||
There were people on HN who still thought that the API prices were being subsidised. The level of conspiracy theory was off the charts on this topic. You would get these price reductions month over month you would still have people believing in crazy stuff.
amazingamazing 3 hours ago|
Link to definitive price info?
dannyw 5 hours ago||
[dead]
spacebacon 3 hours ago||
[dead]
lightinglabs 5 hours ago||
[dead]
shevy-java 4 hours ago||
The milking games have started. The billionaires want their money back.

Edit: Yes, 80% minus is still milking. Because you empower these greedy mega-corporations. Just look at the RAM prices increase, then you see that the more money you give these hungry dragons, they more they will eat up. Don't get fooled by their "less cost now" advertisement.

measurablefunc 6 hours ago|
Model segmentation & distillation like this that asks the consumers to pick exactly which version of the algorithm will solve their problem is evidence for lack of intelligence instead of its presence.
beering 5 hours ago||
You really really don’t need to pick. Just use Sol on high. That’s my daily driver and I don’t touch the model picker at all.

Now, if cost is your concern, then that’s a problem in all of computing. Hence why I’m sending you short plain text messages using an iPhone with a many-core CPU and gigabytes of RAM.

dominotw 6 hours ago||
it is really hard to know upfront if you have fuzzy task. sometimes i would choose a cheaper model and it will spin and spin with bad outputs ending up costing more had i chosen a more capable model.
cute_boi 5 hours ago||
there is mixture of experts which is also another routing. So, simple change in prompt can be a big difference.
More comments...