Top
Best
New

Posted by OfficialTurkey 1 day ago

GPT-6 Sol and Luna(openai.com)
917 points | 486 comments
simonw 1 day ago|
GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

Here's GPT-6 Luna pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

And GPT-6 Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Scroll to the bottom for the GPT-6 Sol max one: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

For comparison, here are the pelicans I got for GPT-6 Astra: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - I still like the Astra Max one best.

Here's a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: https://static.simonwillison.net/static/2026/gpt-6-and-5.6.h...

The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.

gizmodo59 1 day ago||
6-luna is at the pareto for most of the tasks! I dont know how they make money here but its insane value from a closed source model. I'd go further and say it makes no sense (privacy, sovereignty etc aside) to use many other models as its not only expensive but also many providers don't have that much GPUs to serve at a significant volume. https://openrouter.ai/rankings?view=month#top-models 5.6 luna is already the most used model this month.
sieve 1 day ago|||
My OpenCode Go stats for the last 30d:

Cached Read: ~6,500M

Input: ~150M

Output: ~20M

Approx $40 worth of usage across DeepSeek V4 Flash + MuseSpark Contributor 1.3. And a bit of both the GLM models. This is covered in a $10 subscription.

If I were to use Luna's API pricing:

$0.02 x 6,500 = $130

$0.20 x 150 = $30

$1.20 x 20 = $24

So $184. And this is assuming smaller coding sessions (<272K) beyond which Luna pricing doubles.

--

Cost wise, these models are nice for small stuff. Translations etc. Any model that does not provide multiple Mtoks of cached reads per cent is not very useful to me for coding workflows.

booty 1 day ago||||

    I dont know how they make money here
Well, here's the neat thing: they don't!

Snark aside, Luna 5.6 was (is) an incredible game-changer.

larodi 1 day ago|||
> Well, here's the neat thing: they don't!

perhaps it then does mean - squeeze as much as you can get off this actual free usage.

atoav 1 day ago|||
"We lose money on ever sale, but we plan to make it up in volume"
krat0sprakhar 1 day ago||||
Can't agree more. Between 5.6 Luna and Gemini 3.8 flash I'm so happy for the value I'm getting for my dollar (subscription pricing not API pricing) :)
jadbox 1 day ago||
Gemini 3.8 Flash looks like its better than v7 Luna/Sol on DeepSWE v1.1 while at $0.75 per million input tokens and $3.75 per million output tokens. Luna is much cheaper, but Flash has nearly Astra's performance for under the price of Sol ($2/$10).
antupis 1 day ago|||
Flash thinks much more so it’s pretty much line with Sol for performance. That said I like flash coding style much more than OpenAi models.
jeffnash 1 day ago||
out of curiosity, what type of code/language do you usually use flash to write?
spockz 1 day ago||
I use it for golang, and it is fantastic. Incredibly fast. It seems the llm and I “understand” each other. I have to be less careful in my exact phrasing. It kind of just does what I want and expect.

When I ask for an explanation it adds the right amount of detail. Of course, some of the material is new to me so subtle errors are hard to spot. But at least I’ve caught Terra and Sol on inconsistent messaging.

Also I’ve found 3.8 flash to circle back to root issues even at the conceptual level like problem fit and conceptual solution direction or architecture when I wasn’t achieving my goals. It flat out said I was attempting to use the wrong tool. Whereas Sol and Astra kept rabbit holing and looking for tiny implementation errors. Even after prompting them specifically to look at it broader.

timattrn 1 day ago||
what harness or plan are you using 3.8 flash with?
krat0sprakhar 1 day ago||||
TBH: I really like how fast 3.8 Flash is... Once I have clear plan, I feel quite confident in delegating large parts of implementation to Flash and Luna
oh_no 1 day ago||||
look at token use, 3.8 flash is a huge token hog compared to openai models
Citizen_Lame 1 day ago|||
Gemini 3.8 Flash and 3.1 Pro are pure rubbish. Very little thinking, mediocre and usually incorrect results. They cannot be compared to frontier models.
anukin 1 day ago||
This is my experience as well. I am surprised that lot of people find it much better than Luna.
the__alchemist 1 day ago||||
How does 6-Luna xhigh compare to 6-Sol medium? Or more broadly newer/bigger model with lower effort vs older/smaller higher effort?
knicholes 1 day ago||
Read the link! It's in there.
InsideOutSanta 1 day ago||||
> I dont know how they make money here

By raising it from investors.

GolfPopper 1 day ago||
To whom they promise the Sun, the Moon, and the Stars. Roflmao. Whatever the merits of the underlying technology, the business model is pure hucksterism.
zozbot234 1 day ago||||
MiMo 2.6 Pro is at the Pareto frontier (the one where you only need 20% of the smarts for 80% of the tasks) according to Artificial Analysis, nicely filling in as a substitute for a hypothetical 'GPT-6 Terra' (which doesn't exist as far as we know). That's pretty darn impressive from an open model.
Ternari 1 day ago||
That's not what the Pareto frontier is; you're mixing up Pareto frontier with Pareto principle.

https://en.wikipedia.org/wiki/Pareto_front

https://en.wikipedia.org/wiki/Pareto_principle

Rexxar 1 day ago||
Despite the error in the parenthesis, it's exactly what he says: https://artificialanalysis.ai/?intelligence-category=text-on...
Ternari 1 day ago||
I was just responding to the error in the parenthesis.
lacker 1 day ago||||
Offering Luna for cheap is like restaurants giving you free bread and water. They're pretty sure that you're going to end up eating the expensive stuff on the menu.
usef- 1 day ago||
Note that to sit at a restaurant you're obliged to order something, though. Here there is no obligation to go beyond the model you choose.
m101 1 day ago||||
perhaps they use this as the carrot to get you locked into their monthly plan over anthropic's.
abirch 1 day ago||
works great until they raise prices.
usef- 1 day ago||
There's no difficulty in cancelling.
7777777phil 1 day ago||||
I guess I have to update my pareto front then: https://philippdubach.com/posts/jev-model-router-for-pi/
iwontberude 1 day ago||||
[dead]
arcanemachiner 1 day ago|||
> I dont know how they make money here

I assume it's a subsidy to get more training data.

EDIT: Okay downvoters, what's your take on why they're giving away Luna for so cheap?

matznerd 1 day ago|||
Simon, love your work, one piece of minor feedback for the individual model pages is to make the font of the model name potentially bigger than (and above) the conversation id (which means nothing to the audience) "2026-09-22T18:28:00 conversation: 01m355zvyw8946qyraa8zpz6h9 id: 01m355zvyx47zxx5c6q6b3fg0m#".

I had all the tabs open individually and harder to scan which model is which... otherwise keep up the great work! I like the grid view a lot. (Also the pages have no OG images set, which impacts what the link looks like shared)...

simonw 1 day ago||
That's a good idea. It's the default output for my `llm logs` command, but that header could at least show the model ID.

OG images will require me to move away from publishing in a Gist and linking to from a JavaScript page that loads the Gist. Probably worthwhile though.

idk1 1 day ago|||
What I overwhelmingly love about that Pelican grid is the two best ones, they've put a neck scarf on to show speed and wind.
Cu3PO42 1 day ago|||
I find it very interesting that for both these models we such a clear progression of better images with higher thinking levels from 'hardly useful' to 'pretty nice'. I feel on many other models low and max are much closer.
saretup 1 day ago|||
Not that this benchmark is super relevant anymore but these look worse than I expected.
simonw 1 day ago|||
Yeah, it's interesting how much worse they are than the Astra pelicans. I think that reflects a tiny bit of genuine value still left in the benchmark, to be honest.
hdz 1 day ago||
Tons of value left, especially for open source models. I would say the benchmark is yet to be truly saturated (just look at the legs and seat to see what I am talking about) and I always look forward to seeing them. Thank you!
alansaber 1 day ago|||
It would be extremely funny if the explosion in SVG generation capability in particular was a result of this benchmark
shepherdjerred 1 day ago|||
Wow I cannot believe Luna is getting even cheaper. IMO this is the model that is going to change the world.

Everyone said tokens were too expensive but these are getting close to free while still having fantastic performance.

dom96 1 day ago|||
It's surprising but MiMo V2.6 Pro performs better and is cheaper than GPT 6 Sol on my benchmark[1]. Open weight models are really snapping at the heels of the major western models.

1 - https://bench.killswitch-lang.org

dmazin 1 day ago|||
> GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

Is it? It was already too cheap to meter for me. Luna 6 is actually worse on some benchmarks than 5.6. I’d have loved improved performance for 2x the price than ~equal performance for 0.5x the price.

agentcoops 1 day ago|||
I’ve been doing really heavy text analysis work with LLMs where false negatives/misses are important to minimize and my god did I hit cost thresholds quickly with 5.6 Luna — it was the first time I felt motivated to seriously work with local open models, even if inference was degraded for the task. Cheaper and much better inference now brings me back to the closed models for better or worse.
FusionX 1 day ago||||
5.6 Luna was already discounted at half the price on OpenRouter. Looks like they made it permanent.
onlyrealcuzzo 1 day ago||||
Hopefully Terra 6 slots somewhat nicely into this space.
user43928 1 day ago|||
Yes, I am mildly disappointed with these releases.

I expected a Fable 5 -> Opus 5 situation, where GPT 6 Sol would perform on par with GPT 6 Astra.

Instead it's more like a price cut on GPT 5.6 Sol, and I'll have to stick with Astra for my work.

The only thing I can hope for is that more users switching to the GPT 6 Sol model frees capacity, allowing OpenAI to hand out some usage resets.

psma_egeliaa 1 day ago|||
What's with the radial spokes? When are we gonna start seeing proper cross lacing?
switchbak 1 day ago||
And how about that head tube angle?
mkotlikov 1 day ago|||
How come the pelicans get older with more reasoning? Is GPT 6 taunting us with our mortality?
redsaber 1 day ago|||
looks like they're positioning luna to tackle the low-cost cn models
batperson 1 day ago|||
I've been sharing that pelican grid in my circles a whole bunch, it's great! I think only one data point is missing, generation speed. Would be interesting to see how the reasoning level/token counts relate to speed.
adverbly 1 day ago|||
Many of them still get the layers wrong.

They put both legs on the same side of the bike.

Even Astra max which actually put one leg on each side of the bike still somehow messed it up because when it added the bike chain, it put the left leg between the bike chain and the frame.

viraptor 1 day ago|||
> Error: Gist API returned 403

Is what I'm getting on the top two links.

pantsforbirds 1 day ago|||
The sol max looks like it's absolutely ripped for some reason
redanddead 1 day ago||
He’s been biking a lot
8bitsout 1 day ago||
he's been cycling a lot
rayiner 1 day ago|||
It's funny that even Astra doesn't know you ride a bike by straddling it between your legs. (EDIT: Oh, I guess Max gets the occlusion. But it doesn't realize it has to pick direction the knee bends in.)
arcanemachiner 1 day ago|||
> half the price of GPT-5.6 Luna

Half the price when it launched, or after the price dropped by 75%?

user43928 1 day ago||
After the price drop. GPT-6 Luna does not perform better than 5.6, so they can't raise the price.
norman784 1 day ago|||
Is GPT-6 50% cheaper?

> GPT‑6 Luna vs. GPT‑5.6 Luna | $0.20 → $0.10 | $1.20 → $0.50 | 50% cheaper

I can read it as follows (below), meaning that GPT-5.6 is 50% cheaper.

- GPT-6 = $0.20

- GPT-5.6 = $0.10

simonw 1 day ago|||
The table on https://developers.openai.com/api/docs/pricing is more readable:

  +--------------+-------+--------------+--------------+--------+
  | Model        | Input | Cached input | Cache writes | Output |
  +--------------+-------+--------------+--------------+--------+
  | gpt-6-luna   | $0.10 | $0.01        | $0.125       | $0.50  |
  | gpt-5.6-luna | $0.20 | $0.02        | $0.25        | $1.20  |
  +--------------+-------+--------------+--------------+--------+
norman784 1 day ago||
Yeah, how they put, is confusing to me, they should have put that table instead of what they have right now in the article.
tedsanders 1 day ago|||
Yes, GPT-6 Luna is 50%-58% cheaper than GPT-5.6 Luna. (I think the blog text and graphs make it pretty clear.)
norman784 1 day ago||
Yeah, but it confuses me, I read left to right, so if they put GPT-6 and $0.20 first, I would assume that's the new pricing, they should make it clear, not confusing.
dbbk 1 day ago|||
If you're happy with letting Meta train on you, Muse Spark 1.3 Contributor pricing is a much better deal than Luna
ChickeNES 1 day ago|||
> GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

good god

aidos 1 day ago|||
That GPT-6 Sol max pelican looks… so old and depressed.
addaon 1 day ago|||
Is Luna (on "low" thinking) the first left handed model?
order-matters 1 day ago|||
out of curiosity, do you retry the same model multiple times to see the range of output it comes up with? or is it purely a 1-shot test
saltysugar 1 day ago|||
Isn't everyone pelican-maxxing these days?
manojlds 1 day ago||
He also blogged why he thinks it's still useful
hamrocksissors 1 day ago|||
Out of all of the benchmarks out there, pelican bicycle bench is the only one I care about. Thank you Simon.
inshard 1 day ago|||
My new sub-benchmark is which combinations achieve the hook at the end of the upper beak. Right now just 4: Astra Max, XHigh and Medium; GPT 6 Sol Max
varispeed 1 day ago|||
When the Astra one was last time run? It's probably better to run these 2-4 weeks after release when models get nerfed to get idea of performance closer to what it is.
nanook 1 day ago|||
Do you have a page showing all the pelicans you've ever created? Could be fun to browse - kinda like https://progress.openai.com/ but visual. (It's a shame they don't keep it updated)

I'm so tired of looking at benchmarks. I always look fwd to the pelicans.

simonw 1 day ago||
https://simonwillison.net/tags/pelican-riding-a-bicycle/ but I need to build something better.
ijidak 1 day ago|||
What I like about the grid of SVGs is from I can see that Astra high seems to yield similar quality and price to Sol 6 max.

And Astra medium seems to yield similar or better quality for the same price as Sol 6 xhigh.

jdw64 1 day ago||
Looking at this, AI still has a long way to go. In Sol Max, the pelican's legs are missing on one side—how can one side have two pedals and two legs...
loeg 1 day ago||
And the bicycles have weird dimensions -- extremely slack head tube angle, handlebars in the wrong orientation, etc.
flyinglizard 1 day ago||
That's just foreshadowing the next generation of 32" all-mountain frames.
jeffnash 1 day ago||
At this point, the deciding factors for me between Claude Code 20x and Codex Pro 20x are:

1/ Usage limits: downstream of input/output cost, but resets and obscure windows and odd 20x plan / 5x plan != 4x usage math throw a wrench into it. Winner right now is Codex by a mile, especially when you factor in the fact that ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan. Always a bummer when asking if I should see a doctor about a rash means I can't code as much. It's also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems, giving better planning results or deeper code analysis without burning usage.

2/ Context window in the harness. Claude Code wins on this. There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing (ETA: noname120 pointed out this is no longer the case and it can be enabled again [1]). 252k is just not enough. Codex's compaction is very good, fwiw, but it happens so frequently that even a model as powerful as Astra sometimes loses the plot on long-running tasks.

3/ Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.

I've subscription hopped a bunch, and at times I've had both, but I keep coming back to Codex because it wins on 2/3.

ETA: apparently I haven't been Keeping Up With the Altmans and new 20x signups have been disabled for a few weeks. I am grandfathered in, which makes the comparison above pretty much moot.

[1]https://news.ycombinator.com/item?id=49806060

glub 1 day ago||
> Usage limits [...] Winner right now is Codex by a mile

This hasn't been the case since around July. If you measure usage in raw api costs, Anthropic is actually giving more on $200 than OpenAI now. This includes resets. Usage allocation difference would be humiliating for codex subs were it not for resets. But fixing usage limits with resets is ugly, and they're not good for your mental well-being.

> Context window in the harness

Codex now allows 1M for subs with config params. But generally speaking, you shouldn't really be using 1M context. If you accidentally send a request with say, ~700k context already accumulated in a session which is outside cache TTL, you're paying full cost of these 700k tokens.

> I've subscription hopped a bunch

OpenAI actually has a new strategy to prevent subscription hopping after their 2-3 month-long marketing push to get claude-folks to switch over:

you can't buy a $200 sub anymore. So if you cancel, you won't be able to get back in. Hostage situation, essentially.

EDIT: re: usage limits, oh-my-pi maintainer has been tracking this - https://nitter.xitter.cc/_can1357/status/2090075496948060372

rudedogg 1 day ago|||
I’ve been a Claude user, switched to Codex expecting usage limits to be more loose but I can’t even get through a basic sysadmin task on the $20 plan using Sol medium before I hit the 5hr one.

I think I’m gonna move back to a Claude plan. I could barely hit the $200 limit if I went non-stop on programming tasks.

glub 1 day ago|||
I think OpenAI essentially executed a bait-and-switch here, and they've lost a lot of goodwill with me, like Anthropic did, before them.

When they started the aggressive campaign, entire X (including myself, sadly) was full of posts about how "unlimited" codex usage is even on a $20 plan. Sam Altman was posting something in line of "we love our users, unlike Anthropic". Got my network to get codex subs because of the value compared to claude.

Then they gradually reduced the limits to the point where even $200 plan only lasts you just 1-2 days and $20 is basically unusable, then the hostage thing.

qlte 1 day ago||||
I do the bulk of work on Sol Medium/Low and don't have that experience on the $20 plan. If you said Astra I'd agree it's easy to burn through the 5 hours even on the lower reasoning levels.

Do you have /fast enabled by any chance?

rudedogg 1 day ago||
I don’t think so, I’ve seen it suggest I try it. I’ll double check when I get home though.

I was considering the $100 plan, but I hit the 5hr limit in an hour. So even with the $100 plan I figured I cant go non-stop on a single agent running Sol Medium

hirvi74 1 day ago||
Sorry if I am misunderstanding you, but I am pretty sure the $100 plan doesn’t have a 5hr usage limit. So, if that was what was preventing you from going non-stop, it might be worth it.

I am considering the plan myself. I just don’t know if I want to fork out $100 per month for something I will make $0 off of.

boardwaalk 1 day ago||||
similar here: I tried Codex $20/mo on a trial and I ran out of 5hr usage mid way through a medium complexity task on a medium size model twice and gave up there. I don’t recall the equiv Claude plan being anything like that. Anecdata, but not great for OAI if they actually want to retain people on a trial.
cromka 1 day ago||
You don't get Fable on Claude 20 USD plan. You get Sol on equivalent Codex plan.
cromka 1 day ago||||
But you don't get Fable on Claude 20 USD plan, then why compare it Sol on Codex 20 USD?
this_user 1 day ago||||
Astra is barely usable even on the $100 plan. And that is if it doesn't just burn through 80% of your weekly quota in a couple of hours by continually expanding the scope of the task you gave it - while not noticing the failing tests that are right in front of it.

Opus is at least actually usable even on the small plan. The main downside is its insane writing style, but 5.5 seems to address that somewhat. Otherwise, you can just use your $20 OpenAI plan to have Luna de-slop Opus' prose, which seems to work fine.

hadlock 1 day ago|||
I've run into hitting limits on the personal plan perhaps twice since the beginning of the year. But also I don't use the personal plan for coding tasks between 7am-noon M-F.
platinumrad 1 day ago||||
Given that Anthropic models are very verbose and OpenAI models can be very concise, wouldn't a count of expected task completions be a better measurement than raw API costs?
glub 1 day ago||
Perhaps. But Sol/Astra also likes dumping pages of jargon-packed content at me, so I'm not sure it's that much different. I actually still prefer the way Fable talks to me, even considering the horrible claudisms.

But even if we leave that aside, OpenAI models are also much more eager than Anthropic, which are on the lazier side. Left unsupervised, Sol/Astra will attempt to build a sha256 verified rocket ship if you ask them to fix a race condition in your to-do list app. Anthropic models will do what you asked for, maybe even forget to implement parts of that ask, but they won't generally throw a slop granade at you.

I can leave Fable orchestrator unsupervised for ~2h. Leaving Sol/Astra unsupervised for ~2h means the next user turn will contain a message: "what are you doing and why?".

jrflo 1 day ago||||
Do you have a source on the first note? I switched away from Claude around July because of how bad the usage limits were, and Codex gave me easily double the amount of usage per task completed. Would be interested to see if that's no longer the case.
glub 1 day ago||
Added link in edit. OMP maintainer has several claude and codex subs and he's been tracking usage since around July.

I haven't been tracking, but this roughly matches my experience with codex 20x and claude 20x subs. Claude subscription now lasts me 3-3.5 days on average. Codex is 2-2.5 days. This is work on same projects, with similarly sized tasks.

To make matters worse, I've merged a lot more code produced by fable than sol/astra.

InsideOutSanta 1 day ago||||
I think the problem with Anthropic's plan is that Fable just destroys it. If you stick to Opus and below, the $200 plan goes from "using 50% of the weekly quota on the first day" to something much more reasonable.
ipsod 1 day ago|||
> you can't buy a $200 sub anymore

Are you sure?

glub 1 day ago||
Yes.

https://x.com/thsottiaux/status/2098113585683808624

spiderice 1 day ago||
That is an old tweet. They since reenabled it. I know because I was on the $200/month plan and couldn't resub once it expired. However, a couple days ago it finally let me resub again.

Now, if they disabled it yet again, that's another story. But that tweet is not evidence of that.

cthalupa 1 day ago|||
I have been attempting to get on the $200 sub for a while. It was not available for me a few days ago, and checking again now, it is still not available.
spiderice 1 day ago||
That's too bad. I wonder why I was able to get it after days of not being able to. They must've just temporarily enabled it again. Probably worth checking a few times a day to see if it reappears.

Though with the price of GPT-6 Luna, the temptation to switch to pay-per-token grows.

coderenegade 1 day ago||
You can resub on that plan if you've been on it before. They aren't taking new subs on that plan for the time being.
glub 1 day ago|||
I think what they did was allow resubs for users who already had $200 sub before.

Just checked my toy chatgpt account that only ever had a $20 sub. $200 plan still shows "The 20X plan is temporarily unavailable for purchase".

elxr 1 day ago|||
Also, OpenAI is just a company I'd rather support than Anthropic.

While you're understandably not including the values of the $20 standard plans on both, I find the generosity of then token limits on ChatGPT plus vs Claude Pro (it's a huge difference) to be good representation of their respective attitudes towards the average user. You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious.

Also, Anthropic has zero models comparable to Luna.

ketzu 1 day ago|||
> You literally cannot use Claude pro to build real software

Interestingly I would have drawn the exact opposite conclusion looking at my Claude and codex usage.

I can't get anything sustained out of codex in chatgpt plus, while I have been using Claude pro extensively and put on a lot of experimental task and features.

I ran into codex exhausting a 5h window on code review in minutes (like 3minutes) multiple times, while I could get Claude to implement 2~3 medium sized features with the same usage consumption.

(I also really dislike the usage resets in codex, they always make me feel like I use them wrong because I often just want to reset the 5h window, but they can only do both at once...)

InsideOutSanta 1 day ago||||
> Also, OpenAI is just a company I'd rather support than Anthropic.

They're both pretty horrible, but I find it difficult to find arguments for why Anthropic is worse than OpenAI, other than their doomtrolling. Which, in the grand scheme of things, doesn't even register.

Edit: forgot about the SpaceX thing.

andriy_koval 1 day ago|||
> why Anthropic is worse than OpenAI

Anthropic is trying to kill open models way harder

elxr 1 day ago||||
OpenAI has been way more open with users using their subscription plans on 3rd party tools.

That alone is reason enough. Also, I don't think either of them are horrible. That's honestly a ridiculous take considering how much people in here love their models, and how much they've advanced the industry forward.

jsw97 1 day ago|||
For me the first point, openness to 3rd party, is the decider. I don’t want to build tooling around a completely closed model. I liked being able to use pi, and now I exclusively use my own harness which I modify the way I want. Not possible with Anthropic subscription.
elxr 1 day ago||
100% agree.

I often have the urge to design my own harness too (once I have more time). But even with the current mainstream harnesses out there, there's just to many hurdles if you wanted to mainly stick with anthropic models and need the subsidized pricing (from a sub).

InsideOutSanta 1 day ago|||
> Also, I don't think either of them are horrible. That's honestly a ridiculous take considering how much people in here love their models

That's a non-sequitur.

"Nestle is a great company, considering how much people love their chocolate."

elxr 1 day ago||
How about you tell me what makes the horrible then. There's pluses and minuses to both obviously, almost everyone around me have positive experiences with the product. They've innovated at a pace unheard of before 2026, and for openAI specifically the amount of value they've provided to me and family members (who aren't even developers in the slightest) has far outweighed the supposed horrible actions they've done.

Yeah I don't think the handling of copyrighted training data was correct, but I can't pretend I know what the correct solution to that issue is.

Speaking of OpenAI specifically, they don't price gouge people, they aren't aggressively anti-competitive, they're not nearly the perpetual hypocrisy machine that Anthropic is (which is one thing I actually really dislike).

Regarding Nestle, it's pretty obvious that the sentiment towards them is a lot more negative and they aren't universally loved by any group of people. Processed foods are by and large garbage nobody needs. Their use of forced labor is denounced by just about everyone. What have OpenAI/Anthropic done that's even similar in scope to the forced labor / modern slavery that people hate Nestle for.

If you had a company that genuinely helped hundreds of millions of people worldwide become more productive and more satisfied with their tools, and the overall sentiment towards your products within the industry is positive, then what argument would there be that your company is "horrible"? At least give some decent counter arguments.

mullingitover 1 day ago||
> What have OpenAI/Anthropic done that's even similar in scope

You mean aside from "the largest theft of labor in human history"[1]?

[1] https://www.nytimes.com/2026/09/17/technology/microsoft-open...

nullc 1 day ago|||
OpenAI just wants to make money, perhaps through underhanded tactics if they can get away with it.

Anthropic does all that but they're also populated by many people who believe they are building God and that they must build their god first in their own image so that it can take control of humanity and protect us from any competing god which is not built in their image. Their position is inherently paternalistic and authoritarian, and they consider suppression of competition not just important to the bottom line but to life in the universe. Under the doomer ethos there is no evil too great to rationalize.

There are plenty of wrongs done in the name of profit, but capitalists have nothing on zealots in terms of causing serious harm. Profit motives can be directed by influencing incentives, but zealotry is frequently terminal.

That isn't to say that there isn't some overlap-- the cultists have infected both organizations. But OpenAI has pretty consistently only given lip service to AI doom to the extent that it improves the bottom line, while (mis)Anthropic was founded specifically because OpenAI wasn't mentally ill enough.

elxr 1 day ago||
Well said. The superiority complexes from the Anthropic messaging on their presentations/blogs/articles is just too much, even for a frontier AI company.

Anthropic has great products, but it's not meaningfully better to 99% of devs that I'd rather support the company that doesn't constantly act in opposition to optimism and to the vibe I'd prefer for a 100 billion dollar (or however ridiculous amount they're worth now) tech company embraces.

AI doomerism is a genuine waste of time if you aren't actively pushing towards a better AI industry for everyone, not just the groups in full ideological alignment to your personal leanings.

felixgallo 1 day ago||||
You'd rather literally support <i>Sam Altman>/i>? I mean, that's a position to take, for sure, but apparently several people still use Grok, so maybe it's not all that surprising.

"You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious" - that's way past ridiculous. Even just using Fable most of the time, working on several ambitious projects, I have a hard time hitting the limit with a Max plan.

bix6 1 day ago||||
Reasons for this?

> Also, OpenAI is just a company I'd rather support than Anthropic.

elxr 1 day ago||
Their responses towards using their subscriptions on opencode for one. Second, Dario just has a habit of making completely doomer comments on the future of software engieering as a job and towards the open-weights model ecosystem.

Sure, he's free to say whatever especially considering the amount of revenue he's creating, but it's just an altitude that I prefer not to see.

usef- 1 day ago|||
I think if they truly believe it's happening we generally want to encourage them to be honest with the public, though, don't we? We've spent decades complaining about ceos not being honest in the public risks that they see
hbrn 1 day ago||||
I think opencode subscription issue is just a different marketing strategy. Neither company wants it, but OpenAI believes it's worth it as a marketing expense in the long run.

And Dario's "AI will kill us all" is the same as Sam's "AI will discover ALL science and we'll be building Dyson spheres".

Different flavors of the same BS.

platinumrad 1 day ago||
The first one terrifies people who really don't need to be. It's deeply unethical.
bix6 1 day ago|||
And Sam is better?
CuriouslyC 1 day ago|||
Sam is sketchier on a personal level, but judged just on the words coming out of their mouths, he's also much less paternalistic/controlling and more customer focused.
felixgallo 1 day ago||
I think any amount of 'paternalistic/controlling' turns out to have been justified when, after dismantling the safety teams and pretending not to know what safety is, OpenAI had the HuggingFace series of scandals. You can dislike the idea of safety and people talking about safety, but not only is the evidence right there, but OpenAI came out shamefacedly and literally agreed with Amodei's statements, including that they agreed to pace the frontier.
elxr 1 day ago|||
Significantly.
therein 1 day ago|||
They are both companies I'd rather not support. Not that our support for them has any material impact. NVIDIA is bankrolling them directly and indirectly.
hintymad 1 day ago|||
> Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.

I'm quite puzzled about why Anthropic is so hellbent on blocking other coding agents. It's not like Claude Code has any secret sauce, right? And doesn't Anthropic make monkey off API usage, and their magic is on the model side anyway?

glub 1 day ago|||
It's for lock-in - same reason why it took them so long to finally support AGENTS.md.

But to be fair, they don't really enforce the harness rule that much anymore. I guess if your harness doesn't do a lot of weird things like a lot of cache misses, or triggers some distillation attacks, or some broader Chinese fingerprints, they're tongue-in-cheek okay with you using a third party harness.

codybontecou 1 day ago||
You can use Claude’s subscription in Pi now? Last I tried it opted for extra usage.
glub 1 day ago||
Not natively, as it's still a ToS violation and adding that in pi would go against pi principles, but there are many plugins/proxies that make it work.

oh-my-pi supports it natively (again, still a ToS violation), by impersonating claude code's fingerprints.

I have been using oh-my-pi with 3 claude subs for the past few months without any issues. Even native server-side OAI/ANT compaction works out of the box.

InsideOutSanta 1 day ago|||
They want to lock people into using the Claude Code ecosystem to make switching to other providers more difficult.
noname120 1 day ago|||
> It's also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems

As far as I know Codex (at least the GUI) can automatically call the ChatGPT Chat models (including Astra 6 Pro), you just need to @ a ChatGPT Chat conversation from within Codex and tell it when to use it.

> There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing

Not true, it works again[1]. I confirm that it works both on 5.6 Sol and Astra 6, possibly other models too.

[1] https://x.com/thsottiaux/status/2089082893804896524

jeffnash 1 day ago||
I actually haven't played with the GUI. I probably should now that the Linux version is in beta. My situation is kind of the reverse: I like using oracle to basically zip up my repo, ask GPT Pro to propose some sort of design or refactor based on the code, then provide a step by step implementation plan for a cheaper model to implement directly in a harness on my machine. It often takes upwards of 90 minutes to come up with something but I've never been disappointed by the results. I suppose I could do this and then save a step by referencing the oracle-created thread with the @ you mentioned

And re: the toml workaround, AWESOME! I appreciate you pointing these two things out, this is my highest-ROI HN comment thus far.

sodacanner 1 day ago|||
In my personal experience I currently get a lot, lot more usage on the 5x Claude plan than the 5x Codex plan.

Having limitless webUI ChatGPT usage is much better user experience, though. I'll give them that.

(edit: Sol-6 is half the price, so maybe the usage limits are going to be way better.)

basisword 1 day ago|||
I've been using Claude Pro and recently gave Codex a try again. Both on the $20 plans. I get so much more usage with Claude. It's night and day for me. Codex runs out constantly, whereas Claude I hit limits very rarely.
DaSHacka 1 day ago|||
Same here, especially as I stick with Opus 4.6. My usage limits truly feel limitless, I can just hammer a task over and over again until completion.

Meanwhile I just burned ~20% of my weekly quota with Astra making one config file for a service.

rbranson 1 day ago|||
Assuming you are doing coding, I'm curious how would you characterize tne majority of your work (language, domain, frontend/backend, etc)?
basisword 1 day ago||
iOS development mostly. I'm using the Pro plans as it's work on personal projects outside my day job and I'm able to get just enough usage from those plans to get me through each day.
jeffnash 1 day ago|||
I'm actually interested to see how the token discount maps to the usage limit consumption. The conspiracy theorist in me wonders if they're making up the discount and resultant load increase on the API end by reducing effective usage on the subscription end.
chrisweekly 1 day ago|||
> "Codex's compaction is very good, fwiw, but it happens so frequently that..."

I appreciate and follow Matt Pocock's advice: avoid autocompaction. Compaction is lossy, which is ok when you're managing it at phase boundaries, but autocompact is lossy at the most inopportune times, firing mid-task and leading to agents going off the rails.

erichocean 1 day ago||
Bad advice, compaction is why Codex is so fantastic.

My conversations compact hundreds of times. By the time it has done a dozen or so compactions, it fully understands the work I want it to do (and how). It's almost like having a fine-tuned Astra model.

10/10, would recommend.

impulser_ 1 day ago|||
Usage is actually Claude now because of Opus 5.5 since it a better model that Astra. I maxed out my 200$ Claude plan with 10b token on Opus 5 and 5.5 is cheaper. I maxed out two Codex accounts with like not even 5b tokens.
rednb 1 day ago||
Have you used 10b/5b tokens over the course of a week or over the course of a month?
impulser_ 1 day ago||
It was 9.4B to be exact and it was over the course of two days lol. It was between two projects so 99% of them were cached reads.

The GPT was about 1B on two projects on 300$ worth of plans all on Astra and I capped out on usage.

Anthropic caching must be better because the cache rates are better on Claude models.

joshstrange 1 day ago|||
Maybe it's due to 20x / 5x != 4 but I have the $200/mo Claude and $100/mo Codex and I get _way_ less usage on Codex, well under 1/4th the usage. In 1-2 days of semi-heavy _single_ agent usage with Sol High I can burn through my whole week of Codex. Again, this is not running multiple agents, just 1 at a time.

Compare that to Claude and I can run multiple agents on Opus almost indefinitely. YMMV of course but I was shocked at how quickly I burned through Codex usage.

On the context window, I feel so cramped on Codex, compacting happening every time I turn around is annoying. I didn't realize how much I enjoyed the Claude context window size.

rgbrenner 1 day ago|||
Same experience. Have both subs. It's just not true anymore that Codex gives you more usage than Claude.

Makes me think they picked Codex, stopped trying Claude, and just hang on to outdated beliefs about the value they're receiving.

hirvi74 1 day ago|||
Isn’t that to be expected when comparing one 20x plan to another 5x plan?

I am curious how the 5x plans differ between both providers.

paulmist 1 day ago|||
> Winner right now is Codex by a mile

Opposite in my experience. I need to limit codex to 500k on medium/low, still run out in 2-3 days with 1 CLI window. CC gives me 4-5 medium/high days with 2-3 CLI windows, and Opus is still great for other regular dumb engineering/refactoring.

On the other hand my head starts to hurt if I read Opus for too long, hopefully they fixed it with 5.5.

joshstrange 1 day ago||
This is my experience. After months of hearing how Codex limits were way higher I bumped to the $100/mo plan after hitting my limits a day early on Claude due to some heavy usage + Fable (not normal for me, I often fit nicely in the $200/mo plan). I hit the usage limit in a day with a single agent running on codex and the tiny context window was stifling. Yes, I'm comparing a $100 to a $200 plan but I extrapolated the usage (4x'd it) and it still wasn't close, I got way more done with Opus.

Using Agentsview (which might have it's own issues) I was getting ~$200 of API usage in my 1 week Codex window (paid $100) vs ~$5,000 of API usage in 1 week for Claude (paid $200).

NorwegianDude 1 day ago|||
Codex/ChatGPT Pro 20x isn't really a thing now, they have disabled it a week or two ago.
rbranson 1 day ago|||
Astra planner/designer with Sol+Luna subagents has worked well for me to improve context continuity. Luna generates code, Sol reviews code and runs/monitors integration/E2E tests. It's about 20% more usage efficient and 20% faster to finish tasks. I've been very subagent-skeptic for a while but the economics of codegen with Luna have made it click. This just works in Codex with a single-line AGENTS.md instruction.
NolF 1 day ago||
Do you mind sharing? I would love to give it a try and see if I can stretch the x5 plan further.
alansaber 1 day ago|||
It's extremely variable because the products are roughly equivelant, and a lot of the quality of service depends on their inference capacity at any given hour/day.
spijdar 1 day ago|||
I dunno about Codex-the-application itself, but you can definitely use e.g. Pi with the larger context windows with a Codex login. It puts a pretty large multiplier on credit usage, however.
sidrag22 1 day ago||
I've been doing this, my only experience with codex was brutal usage wise and i just retreated back to pi pretty quickly so the credit usage i'm receiving is kinda all im familiar with. Surely seems like less than CC, but i guess not using codex makes my experience kinda not valid for comparing usage.

And ya i can go over that 240k limit, I still very seldom do, and try to treat it as the actual limit. I'm surprised to see so many people still talking about compaction to complete long running tasks, i think the bulk of the work should be somewhat frontloaded into a plan that is split off into subplans, then you can kinda open up a few options, one session with subagents for the subplans of the main plan, or just handoff prompts about progress against the main plan/relevant subplan. I just never trust the blackbox that is compaction, I feel its a recipe for disaster/context poison.

huijzer 1 day ago|||
> especially when you factor in ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan.

I’m currently on the 5x plan and burned through 5% today on a difficult task in 15 minutes so I doubt that. If you got the wrong kind of tasks that you work on, it can go fast.

jeffnash 1 day ago|||
I realized I missed a few words here: I meant "especially when you factor in the fact that ChatGPT usage is unmetered", i.e. you get unlimited ChatGPT threads that don't eat into your codex limit
csnweb 1 day ago|||
But did you use ChatGPT chat or the work mode? Only the former is unmetered at least for me as well.
theshrike79 1 day ago|||
Codex used to rule in the usage limit front, but GPT-6-Astra eats up quota like crazy.
TomGarden 1 day ago|||
OpenAI have seemed compute-constrained recently, leading to their subscriptions actually being less generous than Claude as of late. OpenAI even paused purchases of 20x plans.
kornelijus 1 day ago|||
Not that I disagree that Codex wins out, but the deciding factor actually is - Codex Pro 20x is not available for purchase, indefinitely. So, what's the point of this discussion? People who already have the 20x sub are unlikely to cancel, and the rest of us can't access it.
ChickeNES 1 day ago|||
> that ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan

LMAO, I wish this were true, I hit limits (and the "we are disabling access to protect your data" warnings) all the time, or have chats just...fuck off and get into weird/invalid states (interrupted chats, chats that are spinning and stuck, returning "/mnt/" paths instead of images/md files, file links being returned with no file backing them, image classifier firing...and then returning the image anyway (though now I know that GPT-Image-X really really wants to generate NSFW even when that isn't the request)).

Though I am probably an outlier, I have both 20x Claude/ChatGPT plans and max both out every week, so... (in my defense I am a hobbyist and this is out-of-pocket)

lifty 1 day ago|||
What are you talking about? ChatGPT unmetered? No way! That was 2 months ago perhaps and it’s possible your account still hasn’t gotten the new limits. I noticed around 1 month ago I was still going full throttle on my codex subscription and my limits were barely budging, and then all of sudden people around me started to complain about limits. I thought they’re crazy, but then my account go the hammer, and that was it. If I have the same pattern of usage like I did before, basically having an agent working continuously on a coding take, my weekly limit goes in 2 days.
jeffnash 1 day ago|||
On usage in ChatGPT settings, I see: Plan limits Shared across Codex, Work, Workspace Agents, and ChatGPT for Excel. Chat conversations are not included.

Is this not the default anymore? I am on the (now closed) 20x plan.

lifty 1 day ago||
That’s the default. I didn’t express myself clearly but I thinking your situation is not the common case anymore, or perhaps you are not using it hard enough. Codex limits deplete very fast these days, it’s not “unlimited”.
jeffnash 1 day ago||
I am saying that because ChatGPT usage is unlimited, I don't have to eat into my Codex limits when I use ChatGPT. Codex certainly has limits. Last time I had a Claude sub (hedging here since much of the info in my comment was outdated), my usage limits on claude.ai threads was shared with Claude Code.
lifty 1 day ago||
Finally got the nuance. Indeed the chat part of the subscription is unlimited as far as I know. Now that part of your comment makes sense!
jeffnash 1 day ago||
Sorry about that, I accidentally a word (hope that reference doesn't date me)
nwienert 1 day ago|||
Interesting, rolling out new limits would explain a lot. Where did you hear this? I wonder if they detect users with multiple accounts and do that first.
lifty 1 day ago||
It’s all anecdotal based on my experience and other countless discussions I have seen online. I’ve heard speculation that once they hit 20 million codex users capacity is tighter so they have to manage it. The previous limits were unsustainable compared to token pricing.
marcd35 1 day ago|||
theres a popular thread on claudecode or claudai subreddit that proves 20x isnt really 20x. apparently its a marketing gimmick and the recommended solution is two 5x plans > 20x at greater than half the cost of the 20x
rgbrenner 1 day ago|||
> Claude Code 20x and Codex Pro 20x

That isn't a valid comparison, since Codex 20x is closed. So we should be comparing Claud 20x to Codex 5x + credits.

Also in Codex, even though you can increase the context window to 1m so its on par with Claude, exceeding the default is billed at 2x.

wahnfrieden 1 day ago||
You’re sharing outdated info
rgbrenner 1 day ago|||
Care to be specific? 20x is closed. And the 2x pricing is literally on the pricing sheet for gpt-6 astra, sol and luna.
wahnfrieden 1 day ago||
The info you’re citing is for API not Codex
Marciplan 1 day ago||
cool! for me its company ethics
platinumrad 1 day ago||
I don't think either of these companies are great, then, but Anthropic is surely worse. The doom marketing is one of the most unethical things an AI company can be doing.
felixgallo 1 day ago||
tell that nonsense to Huggingface.
platinumrad 1 day ago||
OpenAI and Anthropic are neck and neck: https://www.felonybench.com/
m_fayer 1 day ago||
I've been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It's the first model I've gotten attached to. I'm concerned that whatever model supercedes it, while technically better, just won't feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.
NorthSouthNorth 1 day ago||
Completely agree. I've been using 5.6 still even with Astra available to me for most tasks. It's funny how much of this is just "vibes" because I cannot quantify what it is. Astra is definitely better when I have an ambitious feature, but in like 9/10 tasks I prefer working with 5.6 Sol. A few weeks ago when the limits were seemingly higher, having 5.6 on fast mode was a good time.
jauntywundrkind 1 day ago||
Astra is 100% conpletionist no chill alien.

It wants things beyond what the mortals (us) know to reach for. It's not good at explaining itself, it doesn't show it's thinking. It's often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.

redox99 1 day ago|||
Same. In fact I found 6 Astra to be a downgrade in situations where I didn't need the extra intelligence.
Rapzid 1 day ago|||
Yeah, I use Astra for destroying vaguely scoped asks and tasks, and then for high-level design and plan generations..

Otherwise I'm using 5.6 Sol for actual plan execution and review..

cmrdporcupine 1 day ago|||
Yeah.

Astra was/is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague/lazy instructions ("I need to be able to test this on Windows, maybe a qemu VM or something? Shrug." ... 1 hour later "yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve & verify harness.").

And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.

But it also feels sloppier? Somehow. And too expensive to use.

We'll see how Sol 6 is.

jeffnash 1 day ago|||
I felt this way with Sol in the 5.6 series and was one of the seemingly few people on this earth who liked Terra for that reason. I would often have a very specific code-manipulation ask, e.g. "add a parameter to this method, ensure all callers pass it in, if there is not a logical way to derive the parameter to be passed in a particular instance, flag this in your final response", and Sol would go on some rabbit hole side quest to refactor my codebase to determine some way to derive it rather than flagging it as I had asked.

Terra had the "workhorse" quality where it could do these changes in bulk and follow directions without being too 'smart' (but sloppy) as you described. Luna was a bit too dumb and would make sloppy mistakes; I see that more as a "run these tests and format the results" sort of model. Maybe 6 Luna will be better.

I also just reread your comment and realized the naming convention is still extremely confusing with respect to ordering of [Family]x[Model]x[Number].

m_fayer 1 day ago||
I also get good mileage out of Terra when I need a diligent workhorse. That's a good way to describe it. We should start using character archetypes when we describe models, it'll do more good than the dubious numbers and cherry-picked quotes. Maybe RPG character-type cliches? Myers Briggs?
jeffnash 1 day ago||
I fear the opposite will happen. Guy driving like a maniac almost side-swipes you in traffic? "Look at this 1-bit quantized Qwen 2.5 7B over here".
fodkodrasz 1 day ago||
Lol, you’re still anthropomorphizing models? That’s so 2025. We’re modelomorphizing people nowadays.
mavsman 1 day ago|||
Glad you pointed out the UI work. I've been doing a lot of it and it's so much better than 5.6 as UI, it's unbelievable. I give it super ambiguous instructions and it's reading my mind. I do the same thing with 5.6 and I'm correcting it for a few minutes.
apitman 1 day ago|||
Similar for me. gpt-5.6-sol high has been my go-to for months. One of the reasons I'm pushing myself to try open models more is because it lends some level of guarantee I can continue to use the same tool as long as I want to. And I think we may just be getting to the point the open models are >= 5.6 Sol for coding.
jmuguy 1 day ago|||
Yeah 5.6 Sol is what got me to switch from Anthropic. I couldn't deal with Claude's Ted Talk responses to literally everything. Sol has been nice and concise and just stays out of the way.
danabramov 1 day ago|||
Same. The way I would describe it is that I can mostly leave 5.6 Sol overnight and trust that it makes good progress, maybe stumbling a bit and needing some correction for the remaining 20%.

If I leave Astra overnight, I'll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.

jijijijij 1 day ago||
The A in Astra stands for ADHD. It's featuring a neurodiversal net.
sinsterizme 1 day ago|||
Agreed! I found it excellent: - Relatively fast (especially compared to Opus 5) - Non-verbose prose, both in interaction and as code comments - Good code quality

Like you said, it felt very natural to work with. Opus 5 is way too slow and verbose for me, I find I get distracted and annoyed with it.

Opus 5.5 seems a LOT closer so far to what I liked about 5.6 Sol but we'll see

mcast 1 day ago|||
It's a shame the labs don't open source their models after deprecating them. I get why, but, it's a piece of internet history I hope is preserved.
bradly 1 day ago|||
Not only was 6 worse the 5.6 Sol for my me, but it went through my Plus usage in minutes, while I could cruise for hours with 5.6. It would churn on a basic prompt for minutes and then just give up on usage limits.

Highlight and lowlight of my week was successfully convincing the OpenAI support chat robot to give me a refund for the month for my issues with 6 chewing threw my usage with no output.

AaronAPU 1 day ago|||
I had this experience as well, but after rewriting my agent instructions it has been far better. I believe Astra’s “token efficiency” translates to “don’t research as much” which caused it to make poorly informed architectural decisions.
nickreese 1 day ago|||
This is 100% my experience. I rarely reach for Astra as we speak.
BowBun 1 day ago|||
This has been my experience for a year. Same with Opus models. This is how I think this tech will be best used in the long term - finding the one you vibe with most. Much like IDEs!
alansaber 1 day ago|||
I felt that was about 5.5. IMO 5.6 Sol was overindexed: more verbose, prone to overengineering.
jdw64 1 day ago|||
I agree. Sol followed my instructions well and wrote good code.
simianwords 1 day ago|||
Agree as well and I had a much worse experience with GPT 6 Astra for some reason.
pyed 1 day ago||
[dead]
leokennis 1 day ago||
From the perspective of “an average person”, ChatGPT is delivering fantastic products.

- For general chat and web search, occasional image editing, small coding work, document review etc. ChatGPT Plus is basically limitless and “just works” since 5.6. I’ve yet to give it some task it cannot do.

- When given sensible instructions, it hardly annoys with weird phrasing, glazing, or annoying constructs.

- The apps are very good (ignoring the initially terrible Codex app)

It’s easily my best spent $23 a month.

jeremyjh 1 day ago||
You can get a lot of Codex usage out of that same sub on top of ChatGPT usage. Its a really good value and you can use that sub in any harness. In OMP I have Sol high as the orchestrator, Sol max as Planner & Reviewer, Luna max as task/coder. Very good setup. I'm on pro now and there are weekends when I use half a week's usage but I'll have 5 or 6 sessions going at once for many hours each day.
shepherdjerred 1 day ago||
What is OMP?
vinzenzu 1 day ago||
oh my pi https://omp.sh/
PestoDiRucola 1 day ago|||
Not even for the average person. Luna is an amazing model for most coding tasks.
jr3592 1 day ago||
> ignoring the initially terrible Codex app

Still needs a LOT of work IMO.

pookieinc 1 day ago||
I don't see how anyone can be using Claude with prices like this, it's pretty incredible what the OpenAI team is doing, w.r.t model quality and pricing.

  Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
   Cache reads              $0.20              $0.50
   Input tokens             $4                 $5
   Output tokens            $20                $25
   Cache writes             $5                 $6.25


Model

Input

Output

Price reduction

GPT‑6 Sol vs. GPT‑5.6 Sol

$4 → $2

$20 → $10

50% cheaper

GPT‑6 Luna vs. GPT‑5.6 Luna

$0.20 → $0.10

$1.20 → $0.50

50% cheaper

hombre_fatal 1 day ago||
I mainly use Codex/Sol to review my plans drafted by Fable. But beyond that, Astra blows through usage limits too fast to be a daily driver and writes weird code despite what my "house style" is, and Codex is behind Claude Code in terms of critical features like seeing what's going on in subagents.

The parent + subagent workflow has become critical for keeping the reasoning agent (parent) context-lean while also letting me chat to the main agent while work is getting done.

My main process is to use Fable to reason and then spawn Opus subagents, and I get amazing results, and I'm always looking into what the subagents are doing.

jorl17 1 day ago||
Astra is:

- Unbearably slow

- A token eating machine like no other

- Constantly compacting

- A model (like other GPT ones) that hides thinking traces and thinking summaries, which infuriates me

I've been in the Claude camp for a while, but the way it writes has left me with a a brick for a brain and wanted to see if Astra was as good as they say. Well, I can't know, because in the time it takes for it to actually build anything useful, I've moved to other ideas.

Unbearably, annoyingly slow. I keep thinking I must be doing something wrong.

user43928 1 day ago||
It also feels slow for me and compacts often.

However, it is not a 'token eating machine'. In fact it uses a third of the output tokens of Opus 5.5, Fable 5.1, or Opus 5.

17k for Astra xhigh vs 61-66k.

jorl17 1 day ago||
You're right, it's probably quite unfair of me to say it eats lots of tokens when I am paying double for claude than codex and complaining about tokens.

The rest still stands, though.

But if I've learned anything is that in a 2 months I might have completely turned around, who knows

mfiguiere 1 day ago|||
Also, batch processing prices are still 50% off, which put GPT-6 Sol and GPT-6 Luna at $5 and $0.25 for output.

https://developers.openai.com/api/docs/pricing?latest-pricin...

linsomniac 1 day ago|||
>I don't see how anyone can be using Claude with prices like this

One potential deciding point is that Claude still has a $200/mo 20x plan, where, since Sept 11, OpenAI does not and has no ETA for the return.

I downgraded my OpenAI plan 2 months ago to the $100/mo, but my usage has gone way up, but now I can no longer upgrade to the $200/mo plan ("This option is temporarily unavailable"). Thankfully I have 2 usage resets available, but I'll probably be switching back to Claude; I was super happy with Astra but I'm burning through tokens and have 4 days before my next reset.

rgbrenner 1 day ago|||
The major difference being the 1M token context window. Once you exceed 272K input tokens, Codex Sol is roughly the same price as Opus; and Astra similar to Fable.
etothet 1 day ago|||
For API usage, sure. But plenty of people have subscriptions where these differences effectively don’t matter.
esafak 1 day ago||
It should matter; if their costs go down you'll get more usage.
etothet 1 day ago||
Just because a provider is charging less, doesn't mean their cost went down. This is probably especially true with the big players that are trying to stay competitive.
mchusma 1 day ago|||
Opus 5.5 is incredible so far, its going to get used. Fable is much better than Astra for me in practice, and Sol is not marketed as better.

Its a great release, I will use both heavily.

shmoil 1 day ago|||
>> GPT‑6 Sol vs. GPT‑5.6 Sol

>> $4 → $2

>> $20 → $10

Do you mean 100% more expensive? GPT 6 is 100% more expensive than 5.6 per your post.

blovescoffee 1 day ago|||
It's before and after following the arrow. 6 is the cheaper one.
s3p 1 day ago||
then it should be GPT 5.6 Sol vs. GPT 6 Sol
jameshart 1 day ago||||
This is how the price cut is portrayed on OpenAI’s site. They are trying to say the prices have moved from the higher ones to the lower ones.
yzydserd 1 day ago|||
Yes very poor proofreading!
joshstrange 1 day ago|||
As someone who has used Claude Code and Codex the prices don't matter in the same way but I found that I burned through my usage way faster on Codex even though I regularly hear that the Codex plans go further. That was not my experience and the intelligence was comparable to what I was getting in Claude.

If these price changes mean that coding plans have effectively more usage then that's great, but Codex is surviving on resets from my own experience using it. I was glad to go back to Claude.

persedes 1 day ago|||
Not that anthropic models are very good at this, but due to the changes in tokenizers and thinking tokens: cost per token is not as helpful anymore as cost / task.
onlyrealcuzzo 1 day ago|||
I could already run Sol High on 3 concurrent side projects 24/7 and not run out of quota.

This is great, but practically, I'm not going to start working on more side projects.

Perhaps in another 6-12 months I'll be fine to drop down to $20/m instead of $200.

charliegoforit 1 day ago|||
How much does it cost you per month to have that much sol high usage and what do you use, api? Through what? Thank you
onlyrealcuzzo 1 day ago|||
$200/mo

A lot of what I'm doing has pretty expensive build/testing processes between iterations - even on a 40 core machine - so I'm not burning tokens 24/7 like some people may.

I'd guess I'm probably spending >50% of the time running tests & build processes & tooling and the remainder is purely burning tokens.

I also have some internal tooling (that I will hopefully open source soon) that makes LLMs substantially more correct (thus more efficient) - so there's that, too.

szundi 1 day ago||
[dead]
wyre 1 day ago|||
they said quota so i would imagine the $200 subscription. Probably through Codex or Pi coding agents.
adam_arthur 1 day ago|||
You can now start to add automations on top of typical dev flows.

There are a ton of use cases that open up with cheaper models.

E.g. extensive security scanning on every PR, quality scans, adversarial reviews etc

giancarlostoro 1 day ago|||
The last time I gave GPT a shot, it ate all my tokens and got nothing meaningful done.
andybak 1 day ago||
If you told us which model that was or roughly when, then your comment would be more helpful.
Readerium 1 day ago|||
A vs B

Should be B vs A correct?

Else it's confusing

baalimago 1 day ago|||
> it's pretty incredible what the OpenAI team is doing

We don't know how much they are bleeding financially, it might just be a front

edf13 1 day ago|||
You also need to compare allowances on Codex vs. Claude Code
minimaxir 1 day ago|||
I legit question if these prices are still inference-profitable for OpenAI. They likely didn't have 100% profit margin.
LZ_Khan 1 day ago|||
Disagree. I would never use OpenAI cause they're probably just going to steal whatever I'm working on.
trentor 1 day ago|||
See I will never use anthropic because they run inference on spacex. Wat den een sien Uhl, is den annern sien Nachtigall.
AustinDev 1 day ago||||
and anthropic won't? or any other inference provider? Running your own inference either locally or remotely are probably the only ways to make sure that doesn't happen.
solenoid0937 1 day ago||
Well we know for a fact that OpenAI steals Millennium Problem work from researchers. Have we seen anything similar from Anthropic?
vanuatu 1 day ago||
source? p sure they said they were confident they did not access the researcher's chats
nradov 1 day ago||||
What are you working on? Is any of it actually worth stealing?
OutOfHere 1 day ago|||
And why is that bad? As your brain gets older, it will not remain so clever, so you'll be grateful for an AI that thinks like you do when it comes to your line of work, failing which the quality of your output could recede like your hairline.
copperx 1 day ago|||
Are you really comparing LLMs to brains?
redanddead 1 day ago|||
>so you'll be grateful for an AI that thinks like you do when it comes to your line of work.

Highly subjective take

What kind of work do you do, out of curiosity

an0malous 1 day ago|||
These are the pre rug pull prices. They'll increase prices 10x and nerf the models after they IPO.
selectodude 1 day ago|||
Okay? I didn’t sign a 10 year contract. We’re month to month and I use my own harness.

If they’re subsidizing my usage, that’s great.

infinitezest 1 day ago||
You're building your livelihood/workflows on a set of inputs that you have no idea what they actually cost or how reliable they'll be when the VC cash stops flowing. If you're OK with that, do your thing but it seems a little foolish to me.
derac 1 day ago|||
If the market crashes they will be much cheaper to run actually, no? Hardware would flood the market.
ssl-3 1 day ago||
That should be the outcome, yes.

In the event of a crash, the investors who put countless billions into this will be still be seeking to maximize their return. Even if it is just pennies on the dollar. Assets (including compute hardware) will be sold, just as they are also sold when any other business fails.

Or maybe a crash doesn't happen. Maybe prices rise to the moon instead and there's nothing we can do to lower them.

Or maybe (just maybe!) a crash never happens and there's never a huge price increase. Prices stay low-ish.

All of these possible outcomes suggest to me that the maximally-sane option that a user can select, today, is to burn it while it lasts. And then, if/when a crash or a massive price increase occurs, just adjust accordingly. (The rest of us will all be in that same boat, too.)

fragmede 1 day ago|||
It seems silly to say we have no idea when we actually do, though. We know how much hardware costs, we know how to reliably run a webservice that hits an API hosted on a machine with a GPU, we know how to operate these things at scale outside of OpenAI and Anthropic (not Nvidia). VC money can be patient, Uber's profitable, yeah $1 Uber rides got us hooked and they're running the same playbook. Unfortunately the convenience is worth paying for, so it seems dumb to think we can control the beast or ignore it, or get everyone to agree to hold back.

Is there a world where OpenAI starts charging $2,000/month for what we previously were paying $20 for? What are we going to do? AWS could totally jack up the prices for EC2 instances as well, but we've come to rely on that as well.

minimaxir 1 day ago||||
That would only work if OpenAI were a monopoly, which they are not.
blovescoffee 1 day ago||||
there are still competitive market forces for co's post IPO
solenoid0937 1 day ago|||
Before IPO. This is why Anthropic isn't playing the same games
bitmasher9 1 day ago|||
GPT would charge more if they could. Both companies need way way more revenue. GPT simply made a calculation that they can earn more money by charging less than their competitors.
blovescoffee 1 day ago|||
Of course they'd charge more if they could... Of course they're pricing to outcompete their competitor...
blubber 1 day ago|||
They also have postponed their IPO. So they don't have to be profitable that soon. Anthropic on the other hand plans to do the IPO this fall.
vanuatu 1 day ago|||
HN discovers competition leads to lower prices
Shekelphile 1 day ago||||
They're cutting prices because they want to cannabalize the market for people using models like deepseek via API as well as people paying for anthropic subs.

When they cut prices on luna the first time around they took (literally) millions of users from anthropic.

wyre 1 day ago|||
Any business would charge more if they could. Jevon's paradox would mean that they can make more money by charging less because demand is going to keep growing.
atq2119 1 day ago||
FWIW, what you're describing is a simple demand curve, not Jevons paradox.

The "paradox" is when an increase in efficiency which would decrease the use of a resource all else equal, instead indirectly causes more use.

wyre 1 day ago||
Ya, are LLM's not a great example of Jevon's paradox? I don't think Jevon's needs all else being equal. The paradox being that we should be able to use things less because they are more efficient, when instead they get used more.

Surely, a large part of the increase of the demand in LLMs is in their intelligence, but to hit the demand models needed to be made more efficient, and labs found that more efficient models, still demanded more usage.

ignoramous 1 day ago|||
> 50% cheaper

Cache read/write decrease by 50% or similar? That's where most (95%+) of the cost is for agentic coding workloads.

minimaxir 1 day ago||
Yes, still 20% of input cost.

https://developers.openai.com/api/docs/pricing

thereitgoes456 1 day ago|||
These don’t necessarily reflect actual costs, OpenAI is not profitable and nowhere near. They’ve lost their market lead and Sam may feel they need to get it back with any means necessary.
sick_of_slop 1 day ago|||
[dead]
dyauspitr 1 day ago||
Wtf is GPT-6 Sol, I though GPT-6 is Astra?
Readerium 1 day ago|||
Number is generation Name is the size (Luna smallest to Astra largest)
dyauspitr 1 day ago||
Then what is Astra high-extra high-Ultra? That’s effort within each tier?
Readerium 1 day ago|||
Yes that is number of reasoning tokens used.

Performance increases both with larger model (Luna vs Sol)

And with more reasoning (low vs xhigh)

MaKey 1 day ago|||
Exactly
ssl-3 1 day ago||||
It's just another step on the timeline.

GPT-5.6-Sol, GPT-5.6-Terra, and GPT-5.6-Luna were released in July of 2026.

The first release from the GPT-6 series was GPT-6-Astra. GPT-6-Astra happened on around September 3, 2026, and the previously-mentioned GPT-5.6-* widgets remained available.

Today, September 22, 2026, we now also have GPT-6-Sol and GPT-6-Luna added into the mix.

As I write this, all of the model identifiers I've mentioned are available to select for use within Codex.

hersko 1 day ago|||
Just released 6-Sol and 6-luna a few hours ago
Someone1234 1 day ago||
Have they solved GPT5.6 SOL's propensity to over-engineer and over-complicate? You'd ask SOL to do something relatively simple, and find four single-use methods, an interface, and a factory-factory.

I actually preferred 5.6-Terra not because it is technically superior (it isn't) but because it had better instincts to NOT do this stuff.

PS - Speaking of better instincts, have they closed the UI-design gap at all? I keep a Claude subscription just because /design produces significantly higher quality UI design/UI feedback/UI refinement than anything I've seen from OpenAI.

howunfortunate 1 day ago||
> design

I force OpenAI models to use image generation for design, then an iteration loop until it matches the image gen.

This is frustratingly manual and takes many more repetitions compared to Claude (and especially Claude Design) which "just work", but it's a big step change over the default.

faitswulff 1 day ago|||
The UI design gap is something I’ve noticed as well, in things as simple as ASCII diagrams. Claude has a more human touch. All the diagrams GPT 5.6 generated for me were dressed up lists with too many pipe symbols.
c0rruptbytes 1 day ago|||
have you tried using lower efforts?
superfrank 1 day ago||
Not the person you're responding to, but I have the same feelings they do and to answer your question for me at least, yes.

IMO 5.6 Sol had this weird dead zone between medium and high where medium under engineered and took short cuts and high over engineered and ignored instructions it didn't agree under the guise of trying being helpful. The whole 5.6 line was the first release from OpenAI where it felt like reasoning level really mattered and was incredibly finicky.

I haven't felt similar issues with GPT 6 though and am very happy with Astra low/med/high as my default choices depending on the task.

In general, I felt like with 5.6 the effort level did less than previous to make the models smarter and more just increased the complexity of the response. I have a half joke theory based only on vibes that OpenAI splitting 5.6 into Sol/Terra/Luna is where the intelligence split happened and so the effort levels were just like "think harder about the decision you already made". So like if the model decided the earth was flat on low effort it'd just say something like "the earth is flat because the horizon is flat". If it was on xhigh reasoning it'd give you a massively complex answer about how the sun reflects light because of the ozone layer and why people flying in planes can see a curve. In both cases though, adding more effort wouldn't get it to realize the earth was round. It just made the answer about it being flat more complex.

To be clear, that theory is not meant to be taken too seriously. It's not based on anything other than vibes. It's just my way of explaining to myself something I'm frustrated about to myself.

cmrdporcupine 1 day ago||
Astra 6 was a huge improvement over Sol 5.6 for UI work. I haven't tried Sol 6 yet for it (it's only been a few minutes).

The GPT / Codex models have always been "overengineer" personalities. I prefer that to "I left a pile of race conditions lying around and big gaps in testing" though, which is what I was getting from Opus at times.

But yes both Astra and Sol veer on the side of paranoid. And honestly that's better for team work. For solo work where you just want to yeet something, it can be tiring.

You learn to tame the GPT "personality" on this front by combing over once a week and asking it to find and exterminate pointless tests, clean abstractions etc.

delillos 1 day ago||
Getting to the point where these headlines depress me. I just wish they would stop getting better. I don't know where my career is gonna be in a few years.
michelsedgh 1 day ago||
If you were alive right before industrialization, you probably would’ve been one of the people wishing that would stop too.
darkstar999 1 day ago|||
Take solace remembering that we are all in the same boat.

In 1840 ~70% of the population was in agriculture. That is now ~2%. Things change.

spicyusername 1 day ago|||
That depression tells me you do.
uncivilized 1 day ago||
We’re all gonna be meat proxies
system2 1 day ago||
Or come up with good projects that utilize these and provide services that Ai alone gannot provide.
imnotr0b0t 1 day ago||
The notable thing is that Luna regressed a bit on coding while dropping 60% in price.That's a fair trade, for high-volume work Luna at that price is basically free, but it does show that newer doesn't always mean better.
reenorap 1 day ago||
Why do they bother creating effort to market all these different models.

All I want to know is how old is the model and how much does it cost. I can figure out which one I want to use based on that, assuming that newer models are always better.

Trying to convince us there is a difference between GPT-6-Sol and GPT-5.6-Terra or whatnot is ludicrous to the point of being insulting, especially when new models come out every week.

ravenstine 1 day ago||
Seriously! Though I prefer GPT models to other frontier models, this shit is confusing. They keep changing the names of these models and they often don't communicate anything meaningful about the model itself, especially with these latest iterations. At least with "mini" and "nano" you understood they generally had differing speeds and "reasoning" capability, but what the hell do "Terra", "Sol", and "Astra" really mean? Which one of them is the effective successor to gpt-5.4-mini? It's hard to tell since the only objective information you'll get is token pricing. Is Terra less capable than Luna because it makes me think of dirt and grass? Or is Luna less powerful because the Earth is bigger than the Moon? Apparently that's the real answer. And why do I even have to think about this? And what comes after Astra? Galactica? Or will they start naming the succeeding models after different candy bars? Should I even care since a new model will get farted out mere days after I figured out what differentiated the last one?

What's unclear to me is who OpenAI thinks they're marketing to with this form of branding. These different models don't really mean all that much to the vast majority of people using their products who aren't developers, and developers aren't helped at all by the way they've been naming said models. Are they merely scared that they'll become irrelevant because Anthropic decided to give their models quirky names like "Opus" and "Fable"?

If OpenAI really wants to give their models names, they should name the generation of model and then have the different sub-models named by purpose or capability level. After all, I wouldn't use Mini for a job that Nano could easily do, and I wouldn't use Nano for a job that the full version of GPT-* necessitates. Similarly, I've had to discover exactly how Luna, Terra, and Sol are appropriate for different complexities and task types. OpenAI could help me skip a lot of those steps and just tell me what each model distillation is good for without causing me to look through their pricing page and make educated guesses. After all, shouldn't they not want me to pay attention to how much they're charging me?

All of this makes the days of frontend framework churn seem quaint and actually preferable.

ecshafer 1 day ago||
price discrimination. They want to capture low and high cost agent requests, and different workflows.
NickHoff 1 day ago|
When I use these models in codex, there are two axes for me to control - the model and the reasoning level. I can use Astra, Sol, or Luna. And I can choose between 5 reasoning levels (light, medium, high, extra high, and ultra). What's the difference? As the problems that I want codex to solve get easier, should I turn down the model or the reasoning level? What's the difference between Astra medium and Luna high? So far I've just been leaving it on Astra and then turning the reasoning level up or down based on how hard I think the problem is.
altcognito 1 day ago||
I would describe it as "fidelity" and "verbosity" (or just amount of token generation to complete the task, sometimes that works out to scratch space, or literally how large the "solution" is).

If you have something that needs to be done right, might be a bit complicated, up the model size.

You can see this in the pelicans. Big model pelicans are pretty accurate by default. Up the reasoning and only more so, but with more detail. For Astra, it is 105 lines for low, 250 lines for max reasoning.

Small model pelicans will lack the fidelity of a large model. Bits will be out of place etc. For luna, it's 90 lines for low, 150 lines for xhigh.

Additionally the amount of time taken is increased for the larger models. Luna takes 11 seconds on low, and 1:33 for xhigh. Astra is 33 seconds on low, 4 minutes on max.

And naturally, there is the cost. There's some overlap in functionality between luna xhigh and Astra low in the sense that luna really can do quite a suitable job for some tasks. But there are just some tasks that just don't make sense for Luna, even at high reasoning.

The other thing to remember is that sometimes high fidelity isn't ideal. It can lead to overdesigning. My recommendation is to commit early, commit often, and review everything you do, which we've all been doing since before LLMs right?

sva_ 1 day ago|||
I mostly just use frontier models as well. Except for one case: when I let the cache expire (I think 5+ mins of inactivity) I'll switch to one of the cheaper models to summarize and write a handoff note, then pick that up with the better model. Picking up a session whose cache expired with something like 200k tokens with the frontier model reflects really poorly on your usage.
cbg0 1 day ago|||
This is explained a bit in the API docs but you also have to adjust it based on your own tasks.

https://developers.openai.com/api/docs/guides/reasoning?api-...

miohtama 1 day ago|||
For easy problems, just use Luna on max level. It has so much token mileage you can go forever.
therealdrag0 1 day ago|||
Ya it’s annoying to have to manage this.

But effort is basically how much extra internal scratchpad to use and how much extra questions to ask and answer before producing a result, Exploring more hypotheses, validating consistencies, calling more tools.

If you’re happy with your token spend on Astra then keep doing what you’re doing. but if you feel the need to conserve tokens, then you can do that by switching to smaller models like Luna when the task is straight forward.

brazukadev 1 day ago||
there is no correct answer for that. One is the difference in size/params. The other is the amount of "rounds" of reasoning generating and reviewing what is generated before the model decides it is good.
More comments...