Top
Best
New

Posted by OfficialTurkey 1 day ago

GPT-6 Sol and Luna(openai.com)
1729 points | 821 commentspage 2
Cu3PO42 1 day ago|
Cutting prices by 50% as compared to 5.6 prices is exciting. GPT-6 Luna at $0.10/Mio input tokens and $0.50/Mio output is positively insane.

EDIT: this doesn't say anything about availability on either Azure or AWS. I'm assuming it will show up later, but it would be interesting if it didn't.

whazor 18 hours ago||
It is insane from a consumer point of view. Luna is cheap and smart enough to do many agentic tasks. Cheap enough so that you can put it on a website without auth.
manmal 1 day ago|||
I’d rather keep 5.6 Sol, and get that even more optimized. I’m not sure I’ll like 6 Sol if it’s anything like Astra.
iyonn 1 day ago||
interesting. in my experience astra has been delightful to work with.
yreg 1 day ago|||
Does API price cut translate into higher allowance on the subscription? Do we know?
manmal 1 day ago||
It does, usually. Luna seems like almost infinite on the 20x plan, and that’s reflected in the API price. Isn’t that the case for all providers?
motoboi 1 day ago|||
already at azure foundry and copilot
c0rruptbytes 1 day ago||
they're already on bedrock
jjcm 1 day ago||
More image->html tests comparing Astra/Sol/Luna:

Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...

All 3 were given the same prompt to dynamically light these and to create the designs as a SPA with page transitions.

Astra: https://html.non.io/annui-astra

Sol: https://html.non.io/annui-sol

Luna: https://html.non.io/annui-luna

Luna gets the button wrong, and in the same way Grok/MiMo did. Looking into it more, it's because Luna actually searched my computer for similar builds, found the ones that I did for grok/mimo, and referenced their files. Astra is still the best by a significant margin in my eyes. Far more polish, better page transitions, effects that aren't overcooked and take into account the page. Better contrast.

alentodorov 1 day ago||
love this eval. keep making them.
wonnage 1 day ago||
The thin serifs not being slightly shifted to align weight-wise with the sans serif is triggering my OCD, but yeah Astra is miles ahead here
davidwritesbugs 1 day ago||
I think there must be such a thing as design dyslexia because they all look fine to me shrug
stelonix 1 day ago||
It seems I'm one of the few Terra users since Astra dropped?

When 5.6 dropped I had no weekly limits and I could just drive my work with Sol xhigh and things were great. Once limits were back (and maybe token prices changed iirc) Sol was no longer usable (on Pro or business) unless I was ok with 4 prompts every 5 hours, so I had to switch to Terra medium/high. I've used Luna for some really dumb tasks like moving files, renaming variables and whatever other old-school refactors I've needed.

Then Astra dropped and it just uses so many tokens I've only prompted with it once. Now with GTP-6 Sol/Luna I'm not sure what's being said here but most importantly I'm wondering whether Luna 6 is a good replacement for Terra.

Has any other Terra user tried and knows more or less than answer to this?

NothingAboutAny 17 hours ago||
I started off with Terra at first before reading anything basically just picking "the middle one" after a while of use I didn't really notice a difference between Terra-Medium and Luna-High, the benchmarks since have suggested there's no real reason to use terra because Luna is twice the speed and some fraction of a cost while on xhigh reasoning achieving better results than Terra medium
stelonix 13 hours ago||
I remember trying Luna on high and finding it spent an enormous amount of time compared to Terra, but after your comment I will try it again and see how it performs. If you're correct, everything will change in my usage.
azuanrb 1 day ago||
Terra is in a weird spot for me. I used to run it as my main driver at medium/high, but after Luna's price drop and some experimenting, I switched to Luna xhigh. If I need extra juice, I just use Sol. Intelligence-wise, Luna xhigh is more than good enough for me. Speed is the only downside. Terra/Sol might be similarly intelligent, but they can get things done faster.
yipinwong 1 day ago||
I've been raving about Luna 5.6 as it's dirt cheap, and "intelligent enough". Double quoted.

Now GPT 6 Luna is even cheaper, and more intelligent, there is no going back... to SOL 5.6 for intelligent layer.

dmazin 1 day ago||
Per the benchmarks in the post, Luna 6 is at best a couple points superior to Luna 5.6 and (unless I’m reading it wrong) xhigh has actually degraded in quality?

I was hoping for a serious Luna upgrade. It was already cheap enough. This feels more like a price reduction than an upgrade.

That said, if the new Luna is able to handle ultra mode and subagents v2 in codex cli, then at least that’s a win.

yipinwong 1 day ago||
Benchmark doesn't really show the whole story.

I forgot which model degraded in quality as time went by, but let's try out Luna 6 for a few more days to confirm for upgradability.

tripledry 10 hours ago||
> Benchmark doesn't really show the whole story.

For me it seems like benchmarks are mostly noise, and the rest is based on vibes. Some find newer models annoying, some are amazed.

elcritch 15 hours ago|||
If you put Luna on Max it's still cheaper than Sol, but can achieve similar results. Though slower and with more iterations. Still it barely nudges my subscription usage!
yipinwong 8 hours ago||
ty for the suggestion. I really never used "max/ultra" on Luna, and will give it a try.
wartywhoa23 14 hours ago||
Ah, the ravers are not what they used to be anymore...
yipinwong 8 hours ago|||
Price is a big selling point for a normie like me.
cindyllm 14 hours ago|||
[dead]
XCSme 14 hours ago||
Also, Terra is gone, GPT-6 Luna is smarter than 5.6 Terra and costs *15x* less [0].

[0]: https://aibenchy.com/compare/openai-gpt-5-6-terra-high/opena...

declan_roberts 1 day ago||
I just switched from Claude to openAI. I'm surprised at how much easier it is to talk to. Claude always spoke to me with a suspicious side eye as if I was trying to do something naughty. For example I could not get it to help me get an old abandonware game running (sim tower).
gizmodo59 1 day ago|
yeah I used to love claude! but these days it refuses and responds as if its like a big brother. glad competition exists and for the past few months codex has been significantly better. Even some oss models like glm are good but they dont have enough compute and get capacity constraints
sfkgtbor 1 day ago||
I'm glad both labs noticed and are trying to improve the models communication styles, they were getting closer and closer to meaningless gibberish.
buckwheatmilk 16 hours ago||
Looks like good one this time. Moving from fine tuned gpt-4.1-mini for structured outputs to gpt-5-mini made zero sense just because 5 was reasoning model and there was no way to disable the reasoning and it also did not have support for fine-tuning.

So essentially I was not able to get nowhere close to the accuracy of previous model and it was slower, and more expensive at the same time.

Now gpt-6-luna, has really competitive pricing and offers similar accuracy compared to gpt-4.1-mini fine tuned for my specific task. And fine tuned models are getting deprecated anyways, seems like a good time to move to gpt-6-luna.

yuretz 17 hours ago||
I wonder what % of comments here are from bots.
phba 16 hours ago||
Maybe I'm imagining things, but every HN thread about a new AI model seems to follow the same pattern, has the same arguments and talking points. The only difference is the version numbers of the AI models mentioned.
wartywhoa23 14 hours ago||
No, you're not imagining, you're seeing a spade for spade. There's absolutely a template they keep rewrapping.

P.S. This cindyllm seems to be stalking me whenever I comment against the grain, does anyone else experience this?

I thought it should have been long dead of all the downvotes it gets, but there we go.

cindyllm 14 hours ago||
[dead]
wartywhoa23 14 hours ago||
No less than 80%.
reenorap 1 day ago|
Why do they bother creating effort to market all these different models.

All I want to know is how old is the model and how much does it cost. I can figure out which one I want to use based on that, assuming that newer models are always better.

Trying to convince us there is a difference between GPT-6-Sol and GPT-5.6-Terra or whatnot is ludicrous to the point of being insulting, especially when new models come out every week.

ecshafer 1 day ago||
price discrimination. They want to capture low and high cost agent requests, and different workflows.
ravenstine 1 day ago||
Seriously! Though I prefer GPT models to other frontier models, this shit is confusing. They keep changing the names of these models and they often don't communicate anything meaningful about the model itself, especially with these latest iterations. At least with "mini" and "nano" you understood they generally had differing speeds and "reasoning" capability, but what the hell do "Terra", "Sol", and "Astra" really mean? Which one of them is the effective successor to gpt-5.4-mini? It's hard to tell since the only objective information you'll get is token pricing. Is Terra less capable than Luna because it makes me think of dirt and grass? Or is Luna less powerful because the Earth is bigger than the Moon? Apparently that's the real answer. And why do I even have to think about this? And what comes after Astra? Galactica? Or will they start naming the succeeding models after different candy bars? Should I even care since a new model will get farted out mere days after I figured out what differentiated the last one?

What's unclear to me is who OpenAI thinks they're marketing to with this form of branding. These different models don't really mean all that much to the vast majority of people using their products who aren't developers, and developers aren't helped at all by the way they've been naming said models. Are they merely scared that they'll become irrelevant because Anthropic decided to give their models quirky names like "Opus" and "Fable"?

If OpenAI really wants to give their models names, they should name the generation of model and then have the different sub-models named by purpose or capability level. After all, I wouldn't use Mini for a job that Nano could easily do, and I wouldn't use Nano for a job that the full version of GPT-* necessitates. Similarly, I've had to discover exactly how Luna, Terra, and Sol are appropriate for different complexities and task types. OpenAI could help me skip a lot of those steps and just tell me what each model distillation is good for without causing me to look through their pricing page and make educated guesses. After all, shouldn't they not want me to pay attention to how much they're charging me?

All of this makes the days of frontend framework churn seem quaint and actually preferable.

wyre 1 day ago||
What models are good at is so subjective it isn't OpenAI's place to really say "Use Sol for X and Luna for Y". They are publishing benchmarks so you can figure out how to best utilize each model. I get that it sucks to have to do this yourself, but eventually there will probably be some type of benchmark that help with discovering a model's strengths and weaknesses

The issue that OpenAI had when they had mini and nano models is that ambiguous the differences between those and everyone just used the base model anyway. I have no idea what type of job mini can do that nano couldn't or vis-à-vis.

I do wonder if it would just be better if they were named 6-small, 6, and 6-big?

ravenstine 21 hours ago||
> I have no idea what type of job mini can do that nano couldn't or vis-à-vis.

In my experience, Nano won't reliably handle complex open-ended tasks and is mostly suited for very explicit instruction that it can't screw up. It's no different from how there are some chores you can give to kids and there are other tasks you need at least a teenager for. If the decision tree of the task is very clear and conventional, Nano can be cheaper than giving the task to a relatively overpowered model, especially if it's something where the output is rigidly structured. This makes it well suited for skills that essentially run CLI commands and generate output, especially because it is usually faster. Mini is more like a discount version of the base model, and Nano is the dollar store version. Mini is more of a generalist and a fairly good deal if you have a moderately complex task that is conventional, but can be less conventional that what Nano can handle. I mostly used gpt-5.4-mini this year for my side projects because it's a pretty good generalist while significantly saving on costs. It is, however, somewhat dumber than the base model and more prone to ignore or forget rules you give it. I'd have just used a base model, but the low cost of Mini and Nano made them appealing to me. Maybe I'm a cheapskate, but I have hundreds or possibly thousands more in my pocket than many other users because of that.

This workflow I settled into with Mini and Nano didn't map cleanly on to the current generation of model tiers. With the price of Luna, you'd think it would be a replacement for Nano. In a sense it is, yet I didn't find that Terra became the new Mini. Terra is more powerful, better at explaining its own decisions, yet I've also found it to be relatively stupid while charging me more to use it. On the other hand, Luna with its reasoning set to "high" is what I consider to fill the role of Mini, and is good enough such that I no longer use Mini. Sol and Astra are great, but they're pricey. It could be my own brain and its bad perception, but so far I don't get the point of Terra. Luna succeeded at reverse engineering some abandonware with a very complicated licensing and virtualization scheme, and did so over SSH into a Windows VM with only PowerShell on the other end. Terra did such idiotic crap to my flashcards app that I stopped using it for anything after that.

This is why I find OpenAI's naming unhelpful and kind of pointless. I don't really care about the benchmarks that all these models are commonly run against. They're not that useful, IMO. OpenAI could easily give early access to these models, get a ton of feedback, and provide better insight to customers on how these things behave. Even calling Terra "gpt-5.6-overpriced-cheating-dumbass" would be better than wasting my time and money figuring it out myself. But that wouldn't make OpenAI as much money.

More comments...