Top
Best
New

Posted by jonotime 2 days ago

Why isn't the industry freaking out about DeepSeek 4.1 Flash?(www.dgt.is)
1096 points | 964 commentspage 11
linzhangrun 1 day ago|
There are too many commoditized models to count: GLM5.3 Flash, Kimi K2.8, Mimo V2.6, MiniMax M3.1...
pizza234 1 day ago||
People have been raving since forever about Deepseek, but if one looks at the CoT, it's evident that it's way way stupider than frontier models (there's a reason why it's cheap). It's laughable to compare Deepseek 4.1 with Opus 5.5.

I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).

Local models are also really slow, unless one spends insane amounts of money.

Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).

apitman 1 day ago||
The argument that most people are making isn't that dsv4.1f is better than frontier, but that it's good enough for most tasks, faster, and way cheaper.

> if one looks at the CoT, it's evident that it's way way stupider than frontier models

Frontier models don't show the full CoT

vintermann 1 day ago|||
It's better to look at the results than the CoT. As far as I know, the CoT is censored for US frontier models - it certainly was for Gemini last time I tried. When you get a condensed summary of the CoT omitting all the false leads, incoherent digressions and backtracking, of course it's going to look smarter.

> I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use

You've measured something, but I'm not convinced you've measured what matters, because that's a lot harder than people give it credit for.

pizza234 1 day ago||
DS4F consistently gives poorer results than Opus 5 (speaking of previous generations) and the CoT shows why.

> You've measured something, but I'm not convinced you've measured what matters, because that's a lot harder than people give it credit for.

"What matters" is what matters to you, right? Who, by the way, don't know what "something" is.

Anyway, if you're so sure that DS performs as good as other frontier models, you're entitled to your opinion. For me that's just having low standards.

vintermann 1 day ago||
An older model that didn't yet have obfuscated thoughts? I'm skeptical. I think the non-obfuscated CoT traces I've seen all look similar.

That's right, what matters to me is what matters to me, and the something you've measured I don't know - but that's not a point in your favor.

The worst sin a model can commit in my opinion, is to give an excellent dazzling response to a slightly different assignment than the one you gave it. DeepSeek seems really good at NOT doing this.

But if you ask the model what it expects to be asked, of course you won't have that problem. It could of course be that DS commits this sin, but just happens to expect the tasks I give it.

But I rather think that it's Claude which is good at expecting your tasks - because I have seen all your "high standards" models commit this sin.

computerex 1 day ago||
The COT isn't an end all be all. Research has shown that the COT isn't necessarily what the model is actually thinking.
kristianp 1 day ago||
> shrank the KV cache by roughly 437X

Can't you just say "shrank to 1/437th the size"? It's not that hard.

tengbretson 1 day ago||
I don't know about "freaking out", but I'd say I'm having a good time here with DS 4.1 flash.
wildster 1 day ago||
I like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md
david-gpu 1 day ago|
Don't you run into it sometimes outputting a few Chinese characters, or Cyrillic, for no apparent reason? I fear it writing some nonsense in the code or the terminal. DeepSeek V4.1 Flash doesn't seem to do that.
jacquesm 1 day ago|||
That hasn't happened with GLM 5.3 yet but with DS 4.1 Flash it did happen and it also had a tendency to loop.
HeavenFox 1 day ago|||
To be fair even OpenAI's and, to a lesser extent, Anthropic's models do that sometimes
Iolaum 1 day ago||
One could argue that all this whole "Pacing the Frontier" bullshit is the industry freaking out regarding the danger of open models.

P.S. That's not to mean there aren't dangers regarding AI. I just don't trust the people making money from selling AI to manage those risks ethically instead of "protecting" us from those risks like pimps but with suits and good manners.

aszen 1 day ago||
Because subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out
s0ulf3re 1 day ago||
I’m partially assuming that there’s a bit of burnout.
patchg 1 day ago||
With two big players thinking of IPOs there is a lot of reasons to whistle on by.
andxor 1 day ago|
Because Haiku is more performant and costs less, even at API prices.
More comments...