Top
Best
New

Posted by OfficialTurkey 2 days ago

GPT-6 Sol and Luna(openai.com)
1763 points | 843 commentspage 7
mrcwinn 2 days ago|
GPT-6 has been fantastic to use. I see Opus 5.5 today but honestly it's been such a rough year with Anthropic, and OpenAI's models are so far ahead, it's tough to consider moving back. I also think OpenAI's desktop app is significantly more polished than Claude CoWork.
fHr 2 days ago||
Luna is the goat for real, cost intelligence ratio is insane already and it is enough for most daily computer use.
msp26 2 days ago||
This Luna pricing is obscene man. 5.6 was good enough for so many use cases (data analysis, structured extraction etc).

Incredible.

lionkor 1 day ago||
I urge everyone to compare this announcement with Anthropic's announcements. From the post above:

> On FrontierCode, which evaluates whether coding agents produce changes ready to merge into real codebases, GPT‑6 Sol improves substantially over GPT‑5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.

I continue to appreciate OpenAI's attempt at some honesty here, showing that they are capable enough and have skilled engineers to a point where they can recognize that slop is hated for good reason, and that there is a real issue. Compare this to anthropic, where e.g. in the Opus 5.5 announcement[1] one of the first points on the page is

> One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.

This is the kind of shit that is the very reason why I stick to OpenAI and deepseek. OpenAI is simply more honest and reasonable about their models' capabilities, while delivering models that still have solid value.

Notice how the OpenAI announcement doesn't make use of anecdotes.

[1]: https://www.anthropic.com/claude-opus-5-5

Havoc 2 days ago||
Interesting to see the US frontier shops cutting prices drastically.

I guess the chinese competition spooked them.

nickandbro 2 days ago||
Pricing is insane, can have Luna going after a goal for 10 days and not run into maxing out the limits.
physicallyIllfr 2 days ago|
Why would you do this though, surely these long running /goal tasks just like letting a wild animal out into your code base.

Does anyone care about code quality anymore?

blovescoffee 2 days ago|||
1. you can have luna clean up after itself and improve code 2. you might be doing something like video-editing, cad modeling, artistic direction, pcb routing, etc. that need to run a long time to "converge"
physicallyIllfr 2 days ago||
I prefer to do these things myself and grow my competency.

This will make me more valuable in the future when everyone has lost the ability to do anything on their own.

fragmede 2 days ago|||
Preparing for the zombie apocalypse seems silly to outsiders, but when it actually happens, who's gonna be laughing?
jpadkins 2 days ago|||
Has there ever been an instance in history when this strategy worked? Plato argued that writing things down will make your memory worse, and less skilled as a debater (kind of true!) How are the Luddites doing at textiles? I remember the arguments that using 'high level languages' like C and Pascal will make you not understand machine specific details (kind of true!)

I respect that you want to learn how things are done, that is a great trait. But once you learn how its done, you should use the tools to free up cognitive load for more difficult tasks.

jasbury 2 days ago||||
For me, long-running tasks are not about generating a lot of code. I’m very picky about what my code looks like. But I’ll happily run for long periods of time debugging problems and/or doing testing and validations. Depending on the problem space, this could mean hours of work for each iteration while it attempts to find a working solution
minimaxir 2 days ago||||
Luna is fine. It's not Claude Sonnet 3.5.
physicallyIllfr 2 days ago||
No its not I use llms just as much as the next guy and not even fable can keep a codebase organized on long running unspervised tasks.
thimabi 2 days ago|||
Not all tasks require frontier intelligence. If you’ve got an easy, but tedious workflow, Luna can be quite good at that.
toephu2 2 days ago||
When using max effort, I run into context compaction quite a lot. I haven't seen any increase in context window size at all over the past half year (stuck at 1M) for all the flagship frontier models.

Have the frontier labs stopped trying to increase context window size?

Alifatisk 2 days ago|
Avoid Max effort for longer conversations, keep Luna at Xhigh, that’s enough.
m3kw9 2 days ago||
The new default is 6.0 Sol high. Escalate to Astra-medium. If usage is tight go luna6.0-max
scosman 2 days ago||
Excluding Opus 5.1 from the coding benchmarks is telling. Opus 5 already matches Astra, Opus 5.1 is much better than 5, and 5.5 is much better again.

OpenAI seems really competitive in most areas, and extremely competitive on cost, but still behind on coding.

vinzenzu 2 days ago|
There is no Opus 5.1. Guess you're talking about Fable 5.1
scosman 2 days ago||
hmm, I was talking about Opus 5.1 but apparently it was a real life hallucination!? Time for bed.
beardsciences 2 days ago|
There's no way this wasn't meant to coincide with Anthropic's release today.
jstummbillig 2 days ago||
They hinted this release last week for tuesday already, so if anything it would be Anthropic that tried to make this happen. But I doubt it.
jrflo 2 days ago||
Altman said it was launching last week on twitter, but they pushed it back to this week
More comments...