Top
Best
New

Posted by jonotime 22 hours ago

Why isn't the industry freaking out about DeepSeek 4.1 Flash?(www.dgt.is)
225 points | 196 commentspage 2
user43928 1 hour ago|
Because DeepSeek is not "a month or two" behind as claimed in the article.

These open models still did not beat February's Mythos / Fable 5.

DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.

Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.

It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.

BobbyJo 1 hour ago||
I was thinking about this earlier today and I came to the following question:

If you had a model 10x as capable as the best model out today, but it cost 100x more, would there be a market, and, if so, how big?

I think there would be a market and I think it would be large.

So, I agree.

lifeisloving 12 minutes ago|||
Many people would, and you'll find that they're building crappy webapps where you dont need SoTA. Like seriously who needs these frontier models?

Unless you're doing some extermely difficult post-grad lvl research, you do not need a 100x PhD research assistant, especially not for whatever silly SaaS product most people are building.

There's people at my job that get so much more done than everyone else using Fable/Opus/Astra. and all they use is the fastest cheapest models. I'd say the people who are using sota models for everything are doing it just because they prefer to be lazy.

You simply do not need these frontier models, they outgrew most people's needs 6 months ago, but for some reason people still want to run a 700k rack of gpus full throttle to center a div for them.

ForHackernews 57 minutes ago|||
Doing what? How many jobs involve solving Millennium Prize math challenges?

99% of everything is CRUD LoB apps.

BobbyJo 17 minutes ago||
I am coding CRUD apps with a mix of astra, sol 6.1, fable and opus 5.5. A more capable model would still benefit me imo. Being able to follow high level guidance better, and being able to harness other models for each task would be a big improvement.
lifeisloving 9 minutes ago||
Do you know how what you're doing, or do you find yourself working on things you dont understand and need the best model because it's the only way to push your own capabilities (because you're avoiding learning how to do the thing yourself)?

Not asking to be mean, I just genuinely dont know why you'd need the frontier for basic applications.

aleqs 39 minutes ago|||
that just sounds like openai/anthropic cope/propaganda, based on absolutely nothing objective lol

even their harnesses are far surpassed by pi and opencode at this point

also sick 'rumors' lmao, apparently marketing through rumors is in vogue these days

enraged_camel 29 minutes ago||
>> that just sounds like openai/anthropic cope/propaganda, based on absolutely nothing objective lol

Nah. There are benchmarks. They are free to look at. And they paint a very clear picture.

aleqs 5 minutes ago||
Yeah the picture they paint is that they're mostly bullshit
ForHackernews 58 minutes ago||
There absolutely is "good enough" and I agree with this author: DeepSeek 4.1 Flash is plenty good enough for all the things I would trust an AI to do at my job.
hmontazeri 2 hours ago||
I had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it
jacquesm 2 hours ago||
If DS4.1 impresses you I would be really interested to see your comparison to GLM 5.3. I switched from the one to the other and even if GLM 5.3 is a bit slower I don't think I'll be going back.
badatnames 1 hour ago|||
DS4 (not 4.1) crossed my dont-care threshold and I genuinely stopped paying attention to new models. I'd love to try GLM 5.3 but I just don't see any point in spending the effort any more. I can get passable intelligence for a bargain price either direct from China or from a ZDR EU provider for a small markup. Paying 10x more will not make me 10x happier, it's unlikely to make me even 1.1x happier now I've got some intuition for the natural limits of these models.

I don't even bother checking how much I spent on API any more, its well under $30 over the past 2 months despite daily constant use. Who even needs a subscription at these numbers?

ctolsen 2 hours ago||||
GLM 5.3 is very impressive and definitely better, but it also at least 4x the price.

On that note I’ve been subbing in MiMo-2.6-pro when cost is an issue, which is super cheap and also performing really well.

pjerem 1 hour ago||||
IDK what happened today but I used GLM-5.3 as usual from Ollama cloud and it was so fast it generated entire documents like instantly.

The reasoning and the result document were done after less than 1 or 2 seconds.

Have Ollama suddenly bought GPU capacity?

LeBit 20 minutes ago|||
There is also GLM 5.3 Flash
pdhborges 2 hours ago||
What inference provider are you using?
arush15june 1 hour ago||
I am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it's absolutely amazing and crazy fast, it's alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5/sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.

I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.

Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.

And it never says no for cyber tasks so that's a big win

pimeys 16 minutes ago|
Yeah. I've been enjoying Coralbricks 250-350 tok/s speeds and it is hard to go back to slower models.

Lithos promises even faster speeds if you want to pay more.

apitman 1 hour ago||
> With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited

My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).

I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.

This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.

youniverse 1 hour ago|
How are you guys doing orchestration? I have been fumbling around in my free time trying to build something for myself but is there a repo or something that just works?
apitman 1 hour ago|||
I might not be the best person to ask. I use Pi harness in tmux and just ask my current agent to spawn interactive pi instances in new tmux windows, create a sentinel file for each of them, and monitor the sentinel files for signals every 2 seconds.

Currently have auto compaction turned off. When the orchestrator's context is getting close to full, I have it write a handoff markdown file and point a fresh agent at it.

I do feel like I'm getting close to the point where I might be ready for something more sophisticated, especially wrt to subagents communicating with the orchestrator.

CharlieDigital 53 minutes ago|||
Many orchestrator harnesses exist.

Check: https://agentmgmt.dev/ and find the one that works for you.

I quite like Paseo (been maining it for a week), but Orca also looks good.

simpaticoder 2 hours ago||
The question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.

The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.

agoodusername63 2 hours ago|
I think it also has a bit to do with the AI sector of tech still moving at lightning speed.

Theres already models that outdo DS 4.1 flash in cost/performance. Luna 6 on max effort for example. Luna also doesn't care what time of the day it is for cost calculation.

And I'm sure by the time people ask why Luna 6 is being slept on there will be another cost/performance king

pimeys 13 minutes ago||
Luna is very slow and bad at agentic tasks. DS runs circles around it and there are US providers providing cheaper rates no matter the time of the day.
wg0 2 hours ago||
While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.

I realized that mistake and guided DeepSeek where it should be.

Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.

sampullman 2 hours ago|
Do you mean Fable 5.1? Or Opus 5.5? I'm not sure what you're working on but for me DS 4.1 flash isn't nearly at their level. For the price it's obvious very impressive, though Luna 6.0 is excellent too.
hirako2000 2 hours ago|||
The problem with benchmarks and proprietary models is that one day a model is best at doing X, another day that's not so sure. And anyway, we are not throwing the same X.

I've found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at hand.

mtrovo 1 hour ago||||
care to share what exactly are you working on?
wg0 2 hours ago|||
Fabble 5.1.
alex-moon 2 hours ago||
I think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.
LeBit 56 minutes ago||
I have subscriptions to OpenAI and Claude but use DeepSeek 4.1 Flash for my coding agents.

It costs pennies and you got really great output.

The author is spot on.

zug_zug 1 hour ago|
I did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.

That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.

0xbadcafebee 7 minutes ago|
[delayed]
More comments...