Posted by jonotime 22 hours ago
These open models still did not beat February's Mythos / Fable 5.
DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.
Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.
It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.
If you had a model 10x as capable as the best model out today, but it cost 100x more, would there be a market, and, if so, how big?
I think there would be a market and I think it would be large.
So, I agree.
Unless you're doing some extermely difficult post-grad lvl research, you do not need a 100x PhD research assistant, especially not for whatever silly SaaS product most people are building.
There's people at my job that get so much more done than everyone else using Fable/Opus/Astra. and all they use is the fastest cheapest models. I'd say the people who are using sota models for everything are doing it just because they prefer to be lazy.
You simply do not need these frontier models, they outgrew most people's needs 6 months ago, but for some reason people still want to run a 700k rack of gpus full throttle to center a div for them.
99% of everything is CRUD LoB apps.
Not asking to be mean, I just genuinely dont know why you'd need the frontier for basic applications.
even their harnesses are far surpassed by pi and opencode at this point
also sick 'rumors' lmao, apparently marketing through rumors is in vogue these days
Nah. There are benchmarks. They are free to look at. And they paint a very clear picture.
I don't even bother checking how much I spent on API any more, its well under $30 over the past 2 months despite daily constant use. Who even needs a subscription at these numbers?
On that note I’ve been subbing in MiMo-2.6-pro when cost is an issue, which is super cheap and also performing really well.
The reasoning and the result document were done after less than 1 or 2 seconds.
Have Ollama suddenly bought GPU capacity?
I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.
Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.
And it never says no for cyber tasks so that's a big win
Lithos promises even faster speeds if you want to pay more.
My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).
I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.
This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.
Currently have auto compaction turned off. When the orchestrator's context is getting close to full, I have it write a handoff markdown file and point a fresh agent at it.
I do feel like I'm getting close to the point where I might be ready for something more sophisticated, especially wrt to subagents communicating with the orchestrator.
Check: https://agentmgmt.dev/ and find the one that works for you.
I quite like Paseo (been maining it for a week), but Orca also looks good.
The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.
Theres already models that outdo DS 4.1 flash in cost/performance. Luna 6 on max effort for example. Luna also doesn't care what time of the day it is for cost calculation.
And I'm sure by the time people ask why Luna 6 is being slept on there will be another cost/performance king
I realized that mistake and guided DeepSeek where it should be.
Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
I've found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at hand.
It costs pennies and you got really great output.
The author is spot on.
That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.