Top
Best
New

Posted by jonotime 2 days ago

Why isn't the industry freaking out about DeepSeek 4.1 Flash?(www.dgt.is)
1093 points | 962 commentspage 9
throwdbaaway 15 hours ago|
Finally an article that gets the maths. Following the release of DSV4 preview where 1M context can fit in single digit GB of VRAM, frontier labs pricing just became stupidly expensive. Other Chinese labs and r/LocalLLaMA also can't compete on pricing.

And since then, there has been so many articles that made it to HN front page, and all of them didn't get it. They just went on and on about tokens generation. Most HN commenters didn't get it either, find-in-page for "cach" typically yield 2~3 responses. If I had a dime for every time this happened, I could have .. paid for 1B cached input tokens?

Anyway, DeepSeek still has to come up with a frontier model, and they almost did it with DSV4 Pro 0813, which is just slightly below GLM 5.3, but 30x cheaper. Unfortunately, the massive price hike happened just 3 days later.

DSV4.1 Flash is good, but not quite the same level. Much easy to self-host though, especially for serving a team of developers. Let's see what the next one can do.

nerdypepper 1 day ago||
https://tangled.org/astrra.space/ds4-recipe is an incredibly cool writeup on making deepseek v4.1 flash run really fast.
wren6991 1 day ago||
It's a solid little model, and I appreciate DeepSeek's commitment to the bit in releasing a brand new pretrain, double the size, numerous architectural innovations as a ".1" release over the excellent DeepSeek V4 Flash.
PaulHoule 20 hours ago||
Because we are still in the phase where we can expect there to be something else to freak out to next week.
rurban 1 day ago||
That's what I thought until August. Used it for half a year almost exclusively. But after the API price increase I'm back at the Claude Pro and Kimi subscriptions.
skc 21 hours ago||
Eventually the party will be over but for now it should be a no brainer first choice tool for some 90% of dev tasks
cyberrock 1 day ago||
Maybe this is a minor issue, but it seems like different providers on OpenRouter etc. have different quant settings. I imagine that that affects the perception of the model quite a bit.
abhinavsharma 1 day ago||
I just keep getting more ambitious with what I use AI for; and that type of work needs surfing the frontier at all times. Because ultimately, many capabilities are not yet saturated
shadyr 1 day ago||
I've been using DeepSeek's API and have been happy with it, but I might look into OpenCode as well. Does OpenCode run a quantised version or use different providers from the official one?
ne01 1 day ago|
Deepseek V4.1 Flash is a hidden gem, really. Not to mention, you can easily get it through many providers that offer zero data retention and consistent speeds above 200 tokens per second!
More comments...