Top
Best
New

Posted by crorella 13 hours ago

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price(openai.com)
866 points | 785 commentspage 12
sehw 13 hours ago|
I stopped using LLMs. I shit you not. My life got better.
Starlevel004 13 hours ago||
Okay, now price cut 6 Sol (and rename it to Terra again).
dools 9 hours ago||
All of a sudden getting competitive on token pricing over the past couple of releases tells me they’re about to kill subscription pricing big time. The subsidised tokens aren’t going to survive the IPOs but if they can capture baseline dev tasks at a cost competitive with open weight models through Luna then capture the frontier token spend as well they could be pretty well placed. The Jarvis bros aren’t going to be able to afford their dashboards though.
ariwilson 10 hours ago||
[dead]
nicolamanzini 10 hours ago||
[dead]
jp1016 11 hours ago||
[dead]
FpUser 7 hours ago||
[dead]
amelius 13 hours ago||
If these models are so smart, can't _they_ select the right model for each task?
skulk 13 hours ago||
the right model for the task is the one that transfers the maximum amount of USD from your pocket to the provider's bank account.
amelius 13 hours ago||
No because then I'll go to the competition.
mholm 13 hours ago|||
Switching models is _very_ expensive in compute (you have to rerun everything from the beginning), and highly variable in cost. Cursor tried doing this for awhile, but inconsistent performance/usage means most users turned it off and pick models specifically.
Mkengin 11 hours ago|||
Github is trying to do that: https://github.blog/ai-and-ml/github-copilot/project-hydrafu...
condour75 13 hours ago|||
I guess the question is, does the Dunning Krueger effect apply to models? The dumb ones might think they're up to the task.
aleph_minus_one 13 hours ago|||
Why don't you simply ask the respective model which model is best for a specific task? :-)
amelius 13 hours ago||
Because it is more work?
SkyBelow 13 hours ago||
These models have a knowledge cutoff that don't just prevent them from knowing about themselves (especially since most data about the model doesn't even exist until after the model is created), but they also don't know about other recent models. Sure, they can search and use other sources, even make some guesses based on the models they do know, but their default stance is more akin to "User asked about model X, model X doesn't exist, maybe it was an hallucination or mistake, let me do a web search...", but that assumes they have web search and are willing to spend tokens on it.

Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.

The speed I'm having to update that document has not gone unnoticed.

godwinson__4-8 13 hours ago|
Let's all boycott and move to Claude until they release 6.1 Astra. I don't like to be teased.

When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.

When do we get the next jump in capability? When is 6.1 Astra released?

ColonelPhantom 13 hours ago|
Isn't Anthropic doing the same, with Opus 5.5 being out while Fable/Mythos is still on 5.1?
wren6991 13 hours ago|||
It's just vibe versioning, right? Fable 5 is a beloved product, it gets a .1 bump to feel close. Opus 5 and Sonnet 5 had a mixed reception, they get a .5 bump to create a sense of distance.

After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.

godwinson__4-8 13 hours ago|||
Is this due to a similar safety concern or just because it's not ready yet for one (or more) of a myriad of possible reasons?

The coverage around 6.1 Astra seems deliberately playing into the dubious, recently headline "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.

Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves back over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.

Without more details on the credibility of the "safety" concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.