Top
Best
New

Posted by tedsanders 3 hours ago

Advancing the price-performance frontier with GPT‑5.6(openai.com)
393 points | 257 comments
GodelNumbering 3 hours ago|
"Half the money I spend on advertising is wasted; the trouble is I don't know which half." -John Wanamaker

This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a very strong model, but separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).

gck1 4 minutes ago||
Luna is comparable to GPT 5.4 from 4 months ago on many benchmarks. I know many who have said during that time, myself included, that if that's the model they had to use for the rest of their lives, they'd be fine.

GPT 5.4 is/was a very capable model.

in_a_society 2 hours ago|||
I don't see why it should be all that difficult. All you have to do is first find a library that implements a decent solution to the halting problem and you're off to the races.
njcornell 1 hour ago||
You can use my script p-noteq-np.sh too if that helps.
fractorial 3 hours ago|||
Cosmically apt username given the substance of this comment.
carimura 2 hours ago|||
Exactly. I haven't reached the "let 1000 agents bloom" mode yet, so currently I'm spending real headspace managing agents doing work, and that work is all important, so why "settle" for sub-frontier models for that work? Maybe I'll get there for non-coding work.
bryanlarsen 2 hours ago|||
Highlighting https://news.ycombinator.com/item?id=49113236 in response.

HN could be run as a BBS on 70's hardware. Instead of using a CPU with ~10 thousand transistors, you're likely using one with ~10 billion to do basically the same thing, and you don't think twice about it.

pimeys 1 hour ago|||
If you have agents and users, you can run evals and see how far the models go. Luna is not greatest in tool calls, but if you define your problem well and the tools well, it is comparable to Gemini 4 Flash with much lower price tag.
satvikpendem 1 hour ago||
Luna is good as an end user model for simple tasks like classification, but not as a coding model. Also do you mean Gemini 3.6 Flash? 4 doesn't exist, and Gemma 4 exists but doesn't have a Flash option.
pimeys 15 minutes ago||
Yeah, been running too many evals in the past week I start to mix the versions up. Probably should sleep...

We run an agent company and outside coding the new Gemini 3.6 Flash and GPT 5.6 Luna are very interesting. Luna can do a bit of research and create reports. Gemini is great for computer use.

For programming it's all Kimi K3 now.

throw2ih020 2 hours ago|||
> separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).

Famously, this is also a problem for human coders in sprint planning.

londons_explore 1 hour ago|||
I get frustrated with a poor quality model leaving my codebase littered with wrong comments, which then later trip up smarter models.
hellojimbo 2 hours ago|||
[dead]
JarJarBeatU 1 hour ago|||
[dead]
odiroot 2 hours ago||
You just need a very strong frontier model to do triage of your tasks.

/s

wmf 2 hours ago||
That's not necessarily a joke; the article proposes exactly that.
preommr 3 hours ago||
> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less,

I don't have the words.

I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.

jpadkins 2 hours ago||
When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon and we will see another 10X improvement in price/performance frontier.

The dynamic GPU clusters will be used for the 5% of tasks, and pushing out the frontier. Also there will be a set of knowledge tasks that are not done today (because they are too difficult for most knowledge workers), that will start being done in the future.

jrflo 2 hours ago|||
Burning the weights into silicon would be many orders of magnitude increase, not just 10x. It's kind of crazy that this hockey stick the AI hype bros talk about seems more and more every day like it might be real
jaggederest 2 hours ago|||
https://taalas.com/ has done it already for a wildly obsolete model. 14000 tokens per second.

https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.

ElijahLynn 1 hour ago|||
Wow! You weren't kidding,

I just tried it too and 14,098 tokens in .05 seconds, I barely blinked and it was done. There was no typing at all appearing on the screen. It just showed up.

https://chatjimmy.ai/chats/01dc66a4-4b1b-4dea-bb5f-926855e37...

WalterGR 47 minutes ago||
That link isn't bringing up your chat, FYI. It just shows the default new chat state.
christophilus 42 minutes ago||||
Wow. This is absolutely wild. I didn't expect that.

If we get to anywhere near this speed for the equivalent of the current models... I don't even know what to think about that future.

iamjackg 2 hours ago|||
Holy crap, I was not prepared for how fast it responded. I just wrote "Just wanted to see how fast you are! Can you write me a quick story about a tiger who lives inside a block of cheese the size of a house?"

I pressed Enter, and the response was instant.

> Generated in 0.037s • 14,205 tok/s

This is unbelievable.

kooi 7 minutes ago|||
"Stochastic gradient descent algorithm in Haskel"

"LMS algorithm in bash"

Just barfed it up lol.

Amazing.

jcul 2 hours ago||||
It's crazy. Are they doing any precomputing as you type, I wonder if you paste a block of text is it the same speed.
nerdsniper 1 hour ago|||
I pasted and instantly hit enter on this prompt: "I generated a filter set using REW v5.31.3 using real-world sweep tone measurements from the room I'm listening in . How can I use it as my MacOS output equalizer so that my spotify music is adjusted for this room and speakers"

and it gave a very reasonable answer in non-perceptible time.

throwuxiytayq 1 hour ago|||
No, it really does take ~0.03s to generate the answer. Try your browser's developer tools and watch the requests.
8cvor6j844qw_d6 2 hours ago||||
I'd like to imagine the things that can be done with this speed and the current frontier models.
throwup238 35 minutes ago|||
Fully interactive games where you can talk to every NPC by text or voice and have an LLM drive the story (with your own meta prompts to guide it, if you so wish). Maybe even have them generate assets on the fly too.

I’m still trying to figure out coding agents. I can’t even begin to imagine the things it would enable. Even the most mundane ideas like LLMs-in-HiFreq-trading have huge implications.

ElijahLynn 1 hour ago||||
Truth, it feels like we're in the dial-up age of LLMs right now. And this Jimmy AI is fiber.
Dig1t 1 hour ago|||
Seriously, if Fable or even Opus was this fast that would be a real game changer.
reducesuffering 12 minutes ago||
RSI will be models better than Fable running faster than this, you won't even need a human in the loop to figure out what to do. The high level goal will be accomplished better than the human in an instant
baal80spam 12 minutes ago||||
Just wow, I made a similar request. Result: Generated in 0.042s • 14,201 tok/s

This is crazy.

HDThoreaun 2 hours ago|||
For what its worth the frontier lab models can surely be a lot faster if they wanted them to be but theyre supply constrained so theyre doing stuff like multi tenancy. Since you cant self host them no one outside the labs really knows speed as a solo tenant
ericd 6 minutes ago||
You can kind of get a sense by running these things at home - I'm currently running Laguna. One interesting thing is that per stream doesn't actually slow down that much with multiple concurrents, because the bottleneck remains the memory bandwidth until pretty significant request depths, and then eventually you hit the GPU's limits. It's one of the big forces that pushes for centralization in this stuff, the fixed costs to run one are huge, the marginal costs of additional tenants, relatively small.
vjvjvjvjghv 35 minutes ago||||
Not an expert on this but wouldn’t this be possible with something similar to an FPGA?
kooi 2 minutes ago||
Weights can be baked into silicone or programmed into hardware ala FPGA, but the context will always be dynamic.

High speed SRAM is where the $$$ is

FuriouslyAdrift 1 hour ago||||
Yep it will be ASICs and DSPs all over again. Orders of magnitude changes.
monkeydust 56 minutes ago||
So which shovels companies are the ones to watch for burnt in silicon models ?
FuriouslyAdrift 43 minutes ago||
Imagine the price of a $9 million NVL72 dropped to about $100, used 4 orders of magnitude less power, was the size of ARM cpu, be bundled with pretty much any electronic device, and ran as fast a frontier AI is today.

That's about how disrupting DSPs were to the industries they arose out of (over a very long time frame).

How would that disrupt the industry?

Yopolo 2 hours ago||||
And don't underestimate how much money Google, Microsoft, Amazon and Meta still have to spend on this tech.

Blocking Fable for sure made it very politicl a lot sooner than i expected it to happen.

and because China already has massive problems of getting access, they are pushing it on hardware too like what Huawai did without EUV.

It seems China is already able to do DUV a lot sooner than others expected.

re-thc 1 hour ago||
> It seems China is already able to do DUV a lot sooner than others expected.

That's the media and in particular US KOLs of all sorts driving the wrong impression of China and other places. China and many other places for example have fast public transport that the US doesn't and can't even imagine today. They're not behind.

China's DUV still isn't that production grade (mass produce-able) so don't get that hyped up the wrong way (in a different direction).

The whole China-is-behind with tech and in particular semi wasn't that they can't. The truth is they spent decades in internal politics and corruption. That all got solved with the bans, so thank the bans! Jensen even said the bans were bad.

coffeebeqn 1 hour ago|||
What does that mean though? Like some kind of a ROM memory ?
bob1029 1 hour ago||
Stacked ROM can, in theory, be a lot denser than anything that depends on a capacitor and refresh cycle.

I don't think it would be that difficult to manufacture compared to other process tech. HBM is really hard to do compared to other memory types.

re-thc 2 hours ago|||
> When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon

Google is already working on a similar idea but more "flexible".

kridsdale1 1 hour ago||
Explain.
jmb99 41 minutes ago||
I'm not who you responded to and I don't have any info on Google. Nor can I explain in detail due to NDAs. But multiple major players are working on something along the lines of what the parent is alluding to.

The "edge" AI landscape (in particular, what you can do with ~5W) is going to be nuts in about 18 months.

captainbland 3 hours ago|||
To be fair we don't really know in terms of prices what's real and what's just investor subsidised attempts at market capture at this point. It could well be OpenAI's attempt to drown Anthropic while they've got the halo product if they feel they've got deeper pockets.
w29UiIm2Xz 3 hours ago|||
Enterprises implemented spending caps and inference providers are lowering prices. Seems they are jockeying for market share.
FuriouslyAdrift 1 hour ago||
Partnerships then consolidation comes next...
minraws 2 hours ago||||
I wouldn't be surprised if they still had some margins since cheaper models are much harder to nail the accurate sizes off, and you still pay 2x for 1M context window.

But if this is even at 400B size it's insanity those inference prices, maybe 10-20% margins, if it's higher I would like to know is it their own chips or maybe they have accurately sized the model to fit on exactly a B300?

Could be a lot of magical things we can only speculate, but from here there likely isn't another 60-70% margin, like I have heard people claim, I would definitely be willing to bet on that.

Could still be a healthy 10-30% margin. Especially with Terra.

platinumrad 3 hours ago|||
We can guess based on the decisions of other inference providers who serve these models.
handfuloflight 2 hours ago|||
Do you mean if other providers will cut their prices in turn?
platinumrad 2 hours ago||
Yes. For example, third-party inference providers serve DeepSeek V4 Flash just as cheaply as DeepSeek themselves, if not even more so. This is very strong evidence that the low price of the model is not subsidized.
hzbdhdjs 2 hours ago|||
[dead]
foobar_______ 3 hours ago|||
Hard to believe numbers. I don't mean that as a critique, but literally I am so impressed. Even if the model is a few percent lower for performance but is 80+% cheaper than competitors and is a US company hosted on US based hyperscaler clouds this is kind of a no brainer. Hard for most businesses to justify otherwise.
rpdillon 2 hours ago|||
This is exactly the model that DeepSeek V4 Flash followed, and it's been insanely successful as a result, even though it's not frontier.
ignoramous 1 hour ago|||
DeepSeek v4 Pro & MiMo v2.5 Pro (Opus 4.6 quality models for code) are insanely cheap for agent-driven work due to their super low cached-input prices ($0.0036/mtok) [0]. For Luna, the cached-input price drop isn't disclosed in TFA, but the pricing page puts it at $0.02/mtok, & that's 5x more expensive.

[0] I am constantly surprised how much work pay-as-you-go with DeepSeek / MiMo will get done. I've barely crossed $2 each in a month of use (~200m tokens).

computerex 1 hour ago||
Absolutely. Although DeepSeek started announcing "Peak valley" pricing which started making me nervous. I have spent $50 usd in July on deepseek and for that much spend I got SO MUCH mileage.

I feel perfectly content in using pay as you go pricing with deepseek. On the other hand, although Anthropic's models used to be my bread and butter for personal work, they are simply too expensive to reach for these days.

ismailmaj 3 hours ago|||
it's 80% less cost, not 80% in efficiency gains, could be that Luna was overpriced to begin with, we don't have much info on the models themselves.

Assuming the efficiency gains are real, I feel like something has to give, maybe worse quality due to aggressive quantization/kv cache compression?

heisgone 2 hours ago|||
Let's suppose each models was subsidized at 70%, so that we only pay 30% of the cost. They would loose much more money per token on the more powerful models. It's in their interest to encourage the use of the less expensive models. Let's say they increase Luna subsidies at 90%. They would still "save" relative to the use of the more expensive models.
mlinsey 2 hours ago||||
High-performing open weight models being released recently, and your customers looking into working with multiple providers as a result, are a great reason to drop prices on your non-frontier offerings.

Although I'm sure there are some efficiency gains, the technology is too new and labs are scrambling to release too quickly to think that the low-hanging optimization fruit has been picked already.

dannyw 2 hours ago||||
Been using OpenAI models since ada/babbage/curie/davinci and at least from my own experience, their APIs feel the same.

If you use Codex it's different, the harness has a lot to do with it and there's definitely been changes including recently.

axus 2 hours ago|||
Something can be overpriced and still lose money.
Yopolo 2 hours ago|||
5-10% over months would still be quite crazy.

But yeah I do'nt want to know what Kimi 3 is pushing buttons inside Anthropic, OpenAI and Google.

Besides any floor: For every year the tokens get faster and cheaper, we will see new things like properly working AI factories which mimic expert teams. A lot more parallism as well.

827a 2 hours ago|||
Vera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.
FuriouslyAdrift 1 hour ago|||
AMD MI400 series is already shipping to customers (basically everybody) and it is crazy fast (8x to 10x faster than the previous gen and beats published Vera numbers in FP8, loses in FP4) and 432 GB per chip. 72 chip unified rack architecture (Helios) already shipping and projected to also beat Vera in NVL72.

MI500 series is supposedly already taping out and they're claiming massive increases (we'll find out end of 2027 prob).

cousinbryce 2 hours ago|||
In a data center that is power constrained but not space constrained they could build out new racks and flip the power from the old racks. Wonder if this will lead to moderately used server GPUs on the secondary market someday.
Yopolo 2 hours ago||
I don't think they overengineered a DC like this.

Besides Nvidia Hardware is still sold out and super expensive. Not a single Nvidia consumer GPU got cheaper at all, Nvidia DGX Spark got more expensive too.

It will be swooped of the market the second it hits the market.

gentlewater 3 hours ago|||
This is gonna put Sonnet 5 in a really awkward spot.
baq 2 hours ago|||
I use sonnet as a smart grep and haiku never and that’s only when I have to use Anthropic at all
heaney-555 2 hours ago||||
Luna is comparable to Haiku, not Sonnet.
827a 2 hours ago|||
Totally untrue. Luna and Sonnet 5 are very comparable: https://artificialanalysis.ai/#intelligence

Luna is an extremely strong model.

re-thc 2 hours ago||
> Luna is an extremely strong model.

By benchmarks, which sadly is a poor measure. Yes Luna is a good model under certain circumstances. Whether it is great for general usage is another story. Sonnet is definitely better when prompts are more vague and it needs to decide things. Luna generally sticks to things very strictly and goes off in bad ways.

mediaman 2 hours ago||
Yes but there is a big, big market for subagents to consume lots of tokens cheaply and condense information up to parent agents. Luna would not be my choice for planning. But an explorer to comb through a codebase to find relevant parts? Or for enterprise retrieval, where it needs to search across many different types of data to see where to focus efforts for a smarter model? Or to wake up periodically to evaluate some conditions and determine if a bigger model should be spun up? Definitely.

I've previously found flash (for all the hate it gets) to be good for these kinds of things. Haiku was fine but it's ancient.

re-thc 2 hours ago||
> Yes but there is a big, big market for subagents to consume lots of tokens cheaply and condense information up to parent agents.

That's again not some "intelligence factor" here. Different agents work for different use cases. Luna wins some. Terra wins some. Sonnet wins some. Flash was really good at exploring.

So I'm not sure what your point is? There's a big market for everything. Even within the market you describe it's likely not a Luna-size fits all either.

Philip-J-Fry 1 hour ago||||
In my real world use Luna is as useful to me as Sonnet. And it gets stuff done faster and follows my instructions more closely.
3836293648 2 hours ago|||
Anthropic basically downgraded all their tiers when they released Mythos, no nah, Sonnet 5 is the successor to Haiku 4.x
bakugo 3 hours ago|||
Sonnet and Haiku were already in an awkward spot, likely by design.

Anthropic's big marketing push this year has been entirely focused on getting people to use Opus via a Claude Code subscription, to the point that Sonnet is almost viewed as the poor man's alternative, and from what I've seen, almost nobody uses it.

Actually, here's an interesting project for all the vibe coders looking for their next front page post: scrape a ton of commits from GitHub with Co-Authored-By: Claude and figure out what the percentage split between Opus/Fable/Sonnet is. I'm willing to bet it's less than 10% Sonnet.

supern0va 2 hours ago|||
>figure out what the percentage split between Opus/Fable/Sonnet is.

This may be misleading, since I suspect many are using a blend through sub-agents. I tend to bias for Fable to orchestrate and Opus for implementation via sub-agents.

StilesCrisis 2 hours ago||||
When I'm paying for it, Sonnet. When work is paying, Opus 5, then Fable if Opus gets confused.
petesergeant 2 hours ago|||
Opus 5 is not strong enough as the top-of-stack model, and feels idiotic after a week or two of heavy Fable usage, to the point where I'm paying for Usage Credits to keep using Fable rather than having to slum it with Opus.
solarkraft 2 hours ago|||
They have no (other) equivalent to nano, so it makes sense that it’s much cheaper now. It may have been better, but it was also hell of a lot more expensive.
WarmWash 2 hours ago|||
Totally possible that humans aren't actually that intelligent.
ceroxylon 2 hours ago||
As well as the existing intelligence being swayed by emotions, hormones, circadian rhythms, stress, peer pressure, propaganda, and survival instincts.
afry1 2 hours ago|||
As if ALL OF THAT doesn't represent inherent and crucial elements of judgement, and therefore"intelligence" itself.

We are not purely rational creatures, thank God. Sometimes those "limiting factors" you listed -- stress, peer pressure, hormones -- are crucial elements of informing the problem solving process and arriving at a decision or a solution that actually works.

All an LLM can do is fulfill a prompt, no matter how misguided, backwards, or incomplete that prompt actually was.

"Go jump off a bridge." Hmm. Dying makes me stressed out. I'm not gonna do that.

customguy 2 hours ago||||
That's a bit like saying a tail is swayed by a dog, as if it could exist without one, or would have anything to do if it did.
subw00f 2 hours ago|||
Why does it matter? This is completely based on data produced by humans.
arjunchint 1 hour ago|||
more like they were facing pressure from chinese models, and dropped prices and now their margins are squeezed
re-thc 2 hours ago|||
> Seeing spikes like this makes me question about where the floor really is.

You mean they increased the price and then cut it back and now it is amazing?

Luna had a price hike vs mini (its previous replacement). The cut now just puts it back in that ball park.

Not that this isn't good news, but what's impressive?

zzleeper 2 hours ago|||
Had to ctrl+f for someone saying this.

I typically do lots of mini calls for research (100s of millions or something in that ball park). Newer models made that absolutely impossible, and the fact that the older ones are starting to get deprecated made me switch to e.g. deepseek for some of my runs. We'll see if I move back after this.

aesthesia 1 hour ago|||
Luna's now cheaper than 5.4-nano (for output tokens). That's a significant improvement.
camel-cdr 2 hours ago|||
this type of thing usually means you are the product
mediaman 2 hours ago||
I don't see how this follows. The cost of nails has fallen by 95% over the last century. It's because the cost of manufacturing has fallen. Not because they are selling the information of nail consumers.

Tokens are not normal software, because they have marginal cost, and I think people who are used to software economics really struggle with this. With token generation there really can be manufacturing cost efficiencies where one producer is just straight up better at serving product at a lower marginal cost.

robocat 10 minutes ago||
> The cost of nails has fallen by 95% over the last century

No it hasn't!

A century ago, some nails cost 2.5% of disposable income, and now the same nails cost 2.3% - only a little cheaper.

The cost of nails has remained remarkably consistent for a century. The problem is that you have ignored the depreciation of money.

Let's assume California prices and income and pick a bigger retail package of nails as you might use for building a house; the numbers: in 1926 a 50lb keg of 4" nails was $2.75 and median after tax income might be $108 per month. In 2026 a 50lb carton of 4" nails is $106 and income might be $4,516. I assume nails are now more readily available and the quality of nails is likely better; and perhaps I should have compared galvinised nail prices.

paytonjjones 2 hours ago|||
According to https://deepswe.datacurve.ai/, Luna at Max at it's previous cost was comparable in both performance and cost to Sol at High.

With an 80% reduction in cost that becomes a ridiculous outlier in efficiency.

visiondude 2 hours ago|||
there is a ton of downward price pressure from Chinese open weight models
Der_Einzige 1 hour ago|||
Until I stop getting downvoted for asserting that these guys are profitable per token, HN is going to continue being pikau face shocked at easily predictable things that any serious AI researcher would tell you, and has been telling you for years now!!!
buckle8017 3 hours ago||
They over purchased hardware.

This is very likely priced below recovering the cost of the hardware but still above operating expenses.

paxys 2 hours ago|||
That’s ridiculous. Every major AI lab is compute constrained. That’s exactly why nvidia is worth trillions today. If OpenAI had a single extra GPU they’d be using it to run another training cycle or experiment for their next model.
infecto 3 hours ago||||
What evidence is there?

I have no idea either way but one thing that detracts from these threads is folks claiming things as a fact without evidence.

qntmfred 2 hours ago|||
sama literally just said they wish they had bought more. the price drops are almost certainly due to good old fashioned hardware innovation (wafer scale with cerebras) and optimizing hardware development based on model architecture and inference costs. other inference providers will try to do the same if they can.

https://www.youtube.com/watch?v=XDB5beon4DY&t=4m20s

pavpanchekha 3 hours ago||
Making Luna, which was already very cheap and extremely capable, 5x cheaper is crazy. I use Sol at work but Luna at home, and while there's definitely a difference, it doesn't feel like night-and-day. After a year of ever-increasing prices it suddenly feels (between this, Kimi K3, GLM 5.2) that prices are falling again.
jedberg 3 hours ago||
> Sol vs Luna

> it doesn't feel like night-and-day.

I see what you did there. :)

deklesen 2 hours ago|||
Good observation! Kudos
oh_no 49 minutes ago|||
I pretty strongly disagree about comparing this to Kimi and GLM, 5.2 was a big price hike for Chinese models, and Kimi K3 was a big price hike to that. K3 was within spitting distance of OpenAI pricing (more expensive than short context Terra, less than long context). And that's after months of OpenAI/Anthropic prices going up.

Now we have an American lab drastically cutting a price, feels like this is the opposite of that trend.

pioneer37 2 hours ago|||
Its just a matter of time at this point.These companies are working day and night to capture the market.
maxdo 3 hours ago|||
is kimi that cheap? it's a very expensive model
pixelesque 2 hours ago||
It's cheaper currently on many of the inference providers.

Personally, I'm having surprisingly good results with DeepSeek 4 Pro at home, which is very good value for money: it's not as good as Claude / GPT 5.6 (I have Co-pilot license at work), but it's still really useful for code reviews, validating thoughts, and especially designing / writing unit tests for new (and old before refactoring) functionality.

And it's very cheap per task. (Flash is even cheaper, but I've had issues with that on more complex tasks where it starts forgetting things and arguing with itself "but wait, let me read the function again").

fy20 1 hour ago|||
DeepSeek V4 Pro is ridiculously priced, especially when you take into account caching. According to the DeepSeek usage panel, 50M tokens have cost me $1.38. It's not the smartest and does like to overthink, but if you have well defined problems it's good for coding. Well... except all your data going to China. I just use it for personal projects.
forsalebypwner 33 minutes ago||
Yup, last month I did ~150mil tokens on DeepSeek v4 Pro for just under $3
mark_l_watson 1 hour ago||||
I toggle back and forth between deepseek v 4 flash/pro on FireWorks.ai using OpenCode. Easy to toggle, I default to flash.
oh_no 48 minutes ago||||
where are you seeing cheap Kimi? pricing I've seen is the same across the board (presumably due to licensing terms) and is in the Terra range.
subarctic 1 hour ago|||
I tried out deepseek v4 pro via a couple providers from openrouter, and it's always getting 429s. Are you running it on your own hardware?
pixelesque 1 hour ago|||
I wish!!

No, I'm using it via OpenRouter in pi.dev - I just used it 30 mins ago... Providers (automatically selected): StreamLake and Baidu Qianfan.

Mashimo 1 hour ago|||
Works fine for me via opencode go.
dominotw 3 hours ago|||
depends on what you are doing. if you are doing verifiable tasks like fixing bugs then any model would do as long as you write the right verification.
simonw 3 hours ago||
> The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%.

If the cost of serving GPT-5.6 just dropped by 20%, does that add up to literally billions of dollars in savings per month?

We know Anthropic spend $1.25 billion renting inference capacity from SpaceX (in two Colossus datacenters) from the SpaceX IPO, but we don't know how much of Anthropic's inference capacity that is (presumably a small fraction, since they were operating on top of AWS and other providers before the SpaceX deal.)

I've not seen any numbers that hint at OpenAI's per-month inference bill, but surely that has to be in the multiple billions of dollars as well.

So 20% is a really, really big deal.

NitpickLawyer 3 hours ago||
~2 years ago gemini2.5 helped write better kernes for itself and (only) reached 1% efficiency gains. Today we're at 20%.
magicalist 56 minutes ago|||
If you optimize program A and manage to wring out a 1% improvement, and I optimize program B and improve performance by 20%, you can see the problem with trying to infer anything from those two numbers.

Edit: searching for the story now, further bolstering the point is that was 1% in training time [1], and the openAI claim is 20% in end to end inference cost. This is a bad comparison.

[1] https://deepmind.google/blog/alphaevolve-a-gemini-powered-co...

NitpickLawyer 35 minutes ago||
> This is a bad comparison.

How so? First, kernel writing (or ML engineering more broadly) is a highly specialised task. Not everyone can do it. It shows that models are getting better and better at (easily verifiable) hard tasks. And you can "hire" that expertise much easier than you can hire the equivalent meatbags. And more importantly you can "fire" them as soon as the task is done. And then hire them 3 months later, when the new model drops. And so on.

Second, 20% gains in inference today gives better end results (i.e. lower overall cost) than 1% in training 2 years ago. Today's models are improving mostly via RL. And RL is highly dependant on fast inference (you want many rollouts for each training scenario). Same for dataset filtering, environment generation, distillation, etc.

dust42 2 hours ago|||
In 2 years from now we will be at 400%. https://xkcd.com/605/ Also, it is called kernels (you have nitpick in your username)
dominotw 3 hours ago||
imagine writing that on your resume

> reduced inference cost by 20 percent saving company x billion dollars per month

paxys 3 hours ago|||
Where are you going to apply to with that resume that’s a step up from your current job though?
petesergeant 2 hours ago|||
The other place, but for more money
bpavuk 2 hours ago|||
lots of places, actually. not everyone wants to be attached to the Silicon Valley culture, and that line alone will guarantee practically any workplace. that person is going to find out what work-life balance is :)
paxys 2 hours ago||
Sure, but those places don’t need such lofty resumes to begin with.
speed_spread 2 hours ago||
With the right kind of credentials, it's not about need, it's about want. Flip the roles and let yourself become an object of desire, an aspirational hire.
tekacs 3 hours ago||||
In this case, and I don't mean this critically, I guess it would technically be, "Instructed model to find efficiencies... reducing inference cost by 20% saving company x billion dollars per month."

I have no doubt that further work was required to enable this, but it's still very cool to be possible to say that.

andai 3 hours ago|||
I think they meant that GPT-5.6-Sol can write that on its resume.
da_grift_shift 3 hours ago|||
Does the model get the credit for its promo packet then? :^)
hirako2000 3 hours ago|||
Contributed to. Can't be some IC who made a few nice PRs
kridsdale1 1 hour ago||
Why not. Jeff Dean and John Carmack exist. Both are L10 SWEs.
bob1029 3 hours ago||
This feels like the dialup->broadband transition to me.

I was already a huge proponent of Luna for things like deep research. Being able to run 5x more for the same cost is simply bananas. We are already running 10 parallel agents for hypothesis generation. I cannot imagine 50. The statistics become much more interesting & powerful when you can run so many samples of the exact same prompt+model without breaking the bank.

Imanari 2 hours ago||
How do you run 'deep research'?
dannyw 2 hours ago|||
Deep research is basically a LLM with web search, and a "work really hard" goal-orientated prompt, and some output formatting suggestions.
kridsdale1 1 hour ago||
And self forking fan out.
cg5280 1 hour ago|||
It's a feature offered in ChatGPT and other platforms, though probably gated behind paid subscriptions.
fy20 50 minutes ago||
Used to be part of the $20/mo plan but it's not anymore (not sure if they removed it completely). However GPT-5.6 is pretty good at researching if you prompt it right, I've regularly had it spend 5+ minutes researching topic with lots of web searches.
kaufmann 4 minutes ago||
They placed it in the Plugins submenu.
jrflo 2 hours ago|||
Very interesting. Can you share more about your hypothesis/research pipeline? I have been using Sol for those types of task because I figured you'd need more reasoning for getting good ideas, but maybe quantity > quality at a certain point?
bob1029 2 hours ago||
Here is a rough approximation of the pipeline I use:

Phase 1 - Run X copies of Luna in parallel over the user's prompt. The purpose is to generate a diverse set of hypotheses.

Phase 2 - Run Y copies of Terra in parallel to investigate the hypothesis results, with each receiving them in a randomized order.

Phase 3 - Run 1 copy of Sol over investigation reports.

The goal is to ensure that the agent covers more initial starting points before presenting a final conclusion. If you only run a single copy of Sol and it hooks onto something wrong, it might not recover.

eevmanu 1 hour ago||
How do you run this phases and parallelization on each?

Via just ... "prompting it"?

Or do you use any tool in the middle to ensure this agent architecture?

Just curious if there is any workflow-like tool in the middle that is helping.

bob1029 1 minute ago||
[delayed]
andai 3 hours ago||
Do you have a sense of which tasks benefit from more agents and which don't?
bob1029 2 hours ago||
Anything related to reading and interpreting the environment seems to always benefit from the addition of more agents to the search party, assuming you have some rational way to synthesize their results.

Taking actions that mutate the environment is a different story. I think this is where you run into diminishing returns very quickly. You generally want one strong agent to act given the results of all the searching that was done. If the plan is clear, you don't need a genius model to execute it.

handfuloflight 2 hours ago||
I definitely think you want the genius model to synthesize everything that rolls up to them.
andai 1 hour ago|||
I think this is an unsolved problem. The most interesting thing I saw here is the Recursive Language Models paper.

https://arxiv.org/abs/2512.24601

There's also a great write up here by the author:

https://alexzhang13.github.io/blog/2025/rlm/

kridsdale1 1 hour ago|||
This is why in your brain you have trillion threads processing and summarizing sensor data (immutable functions), but a SINGLE thread of “execution” which we call the conscious soul.
__jl__ 2 hours ago||
Didn't expect that. Luna pricing is crazy now. I don't think there is anything on the market that competes at this price-performance point.

For our production app, OpenAI clearly is the best provider now. Their API is very reliable and has many nice features. The price-performance of the model lineup is incredible. We used open weights model via Fireworks for a long time (e.g. Kimi K2.5). Fireworks is a great provider but we still ran into issues here and there (Same with Anthropic and Google). OpenAI just works, is fast and in my view has a better price-performance ratio across almost all levels of intelligence.

dannyw 2 hours ago|
OpenAI's APIs are extremely reliable for sure. I don't even remember when the last incident or downtime was.
amluto 2 hours ago||
This doesn’t quite count as “API”, but OpenAI’s roll out of OAuth device code authentication was poor, to say the least.
quirino 3 hours ago||
I generally just check the Price/Performance graph on Openrouter: https://openrouter.ai/rankings#performance#benchmarks. Activate the "Show Pareto" toggle on the right.

I was still using GLM-5.2 in my personal projects, but this just made Luna a very easy choice.

qingcharles 1 hour ago||
Hasn't OpenRouter had Luna and Terra on 50% off sale since they launched? I wonder what will happen to that.
quirino 1 hour ago||
It's still 50% off apparently. Listed as $0.10 for input (original price was $1.0 without this reduction or sale)
hattimaTim 2 hours ago||
The official doc says, Luna = Previous Nano models, kind of. Is it really good at coding?
PhilippGille 2 hours ago|||
Depends on the reasoning effort, see https://deepswe.datacurve.ai (add Luna via model selection drop down, if it's not shown by default)
hattimaTim 2 hours ago||
Thanks for the link!
paxys 2 hours ago||||
Smaller models are great if you are doing targeted changes in existing codebases. Don’t expect to use it for creating complex architecture from scratch or do major refactors. The larger the context, the greater the drop off will be.
quirino 2 hours ago|||
According to the link I mentioned above it's roughly as good as GPT-5.4. Haven't tried it in practice yet.

I bet it must be better in some contexts and worse in others.

tosh 3 hours ago||
80% price cut for luna is a very aggressive pricing move

makes it by far the best choice for most workloads that do not need bleeding edge intelligence (reminder: luna can be comparable to opus 5!)

heaney-555 2 hours ago||
Luna is meant to compete with Haiku. What tasks are you seeing it equal Opus on?
dannyw 1 hour ago|||
You'd be surprised at what Luna can do, especially on xhigh or max. It's capable of working overnight, usually productively, just like Sol.

Haiku 4.5, on the other hand, is comparable to performance to Gemma4 31B (with working tool call formatting) in my experience, and Gemma4 strongly wins on vision and multimodal.

newtwilly 2 hours ago||||
According to the Artificial Analysis benchmark graph in the article, Luna can now outperform Sonnet 5 and Opus 5 low at ~4-10x less cost
euazOn 2 hours ago||||
Per Artificial Analysis:

- Haiku: 30 points

- Luna Medium/High/Xhigh/Max: 38/46/49/51 points

That's a massive difference:

- 30 points is Gemma 4 31B territory

- 50 points is GLM-5.2 (744B) territory.

tosh 2 hours ago|||
luna is way better than haiku 4.5
nateb2022 2 hours ago||
I use Luna a lot (over 1T tokens since it came out) and I'd rank Luna (high/xhigh) on par with Sonnet 5, without hesitation.
ls_stats 1 hour ago||
Is it though? I saw a noticeable difference between Sol (xhigh) and Luna (max). Sol appears to understand better my prompts, you need to be more specific/clear with Luna.
NortySpock 2 hours ago||
> In a compute-constrained world where model demand is growing faster than capacity

I don't buy it.

There have been recent weeks where some of the mid-level models (Hy3, Laguna M.1) are free (true for parts of June and July, see Hy3 in Cyan) . Even then the total token usage appears to be reaching a steady-state.

https://openrouter.ai/rankings#top-models

^ the first graph is tokens per week across all models

I guess we just can only throw ideas at an LLM at a certain rate.

I still have ideas and now I can have an LLM vibe code what I want, but I'm not going to let an agent just run unattended for longer than a few minutes or a few bucks for hobby projects.

So maybe it is a matter of lowering the cost of an LLM so I can let it churn for hours at a cost of pennies... But I suspect demand for tokens is very price-elastic.

Yopolo 1 hour ago||
These are free due to some different type of reasons like Nvidida sponsoring free tokens or the model companies.

My company checks the models and pays for Opus through AWS.

You still send the WHOLE context of whatever you want to do to a random endpoint on the internet. If you want to write a good email, you give that context your email address, names, the reason for it etc.

Big companies don't randomly use some random api endpoint to do so.

Anthropics quarerly revenue is still growing very fast. I don't think we have seen even the real potenzial of it yet at all.

Not only are still a lot of countries missing which do not even use anthropic or any other frontier model yet but also all the agentic based solutions enterprise companies are currently building on mass (at least in my industry)

Der_Einzige 1 hour ago||
Openrouter is not even CLOSE to the majority of tokens. I don't know where this myth came from that they account for even ~5% of total token spend! It's not true!!!!

Stop rejecting what we have been collectively telling you guys! LLM providers are profitable, and have been for awhile!

ninjahawk1 2 hours ago|
80% less for Luna is absolutely crazy, in my opinion we may reach a point in the next year where powerful models on the API could potentially be cheaper than subscriptions. Compute just keeps decreasing in price.
HDThoreaun 2 hours ago|
API will never be cheaper than subs because theres a ton of value created for companies by locking people into subscriptions that tend to be sticky.
More comments...