Top
Best
New

Posted by jonotime 21 hours ago

Why isn't the industry freaking out about DeepSeek 4.1 Flash?(www.dgt.is)
225 points | 196 comments
vishvananda 38 minutes ago|
The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.

I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.

This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.

pimeys 1 minute ago||
Yes. You have to find the provider with pricing that suits your usage.

I am having 98% my input in cache, so using Coralbricks makes sense due to them giving cache reads for free — you only pay for writes. I spend maybe 5-10 dollars a day and my agents basically work day and night implementing things for me.

If your tasks are write-heavy, find a provider with cheaper output.

If you build a customer-facing app, pay a bit extra for 400+ tok/s e.g. on Lithos.

georgel 10 minutes ago|||
I am curious how you managed to spend that much on Deepseek via OpenRouter. I loaded $100 back in July while using v4-flash or whatever the cheap good model was at the time, and have upgraded as the new ones came out from Deepseek. I still have $16 and some of that spend also goes towards the AI usage from my customers (the context they need to load in is quite large too).

And I am using the Claude Code harness with DS as the endpoint. And I use it ~5-8hrs a day to do my coding.

onlyrealcuzzo 10 minutes ago|||
OpenRouter is complete garbage.

Buy directly from DeepSeek's API.

You can literally get overcharged 100x on DeepSeek on OpenRouter (or more).

georgel 7 minutes ago||
I'm all in for saving money and _can_ move to using DS directly from them, but maybe I am missing something here:

OpenRouter Pricing:

$0.02/M input tokens $0.60/M output tokens

DeepSeek Pricing (cache miss, off-peak):

$0.15/M Input $0.60/m output

girvo 2 minutes ago|||
When 98.5% of my requests are cache hits (according to Pi for the last week), the cache miss price isn’t that important to me, and $0.003-0.006 per 1M input tokens is shockingly cheap.

It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?

mswphd 3 minutes ago|||
I've heard that certain inference providers may have different quality of caching implementations, so even if the listed numbers are as you say, the practical cache hit % you get might be significantly different/incur significantly different costs.
lionkor 21 minutes ago|||
How's the caching? I have 99.5% cache hit rate with deepseek when using their own API, it's dirt cheap.
seunosewa 13 minutes ago|||
Which provider was that?
mensetmanusman 36 minutes ago||
Also DeepSeek usage is subsidized as well, it’s a power hungry model.
jchw 33 minutes ago|||
Interesting. Are all of the providers on OpenRouter simply losing money? How does that even work out?
worldsavior 16 minutes ago|||
You're the RLHF.
jansan 21 minutes ago|||
Don't ask and dance as long as the music keeps playing.
kennywinker 32 minutes ago|||
Are you sure about that? My impression was most providers on openrouter were purely selling tokens for profit...
thesnarkitecht 24 minutes ago||
Have y'all tried an Ollama Cloud subscription? Their off-hours pricing for V4.1 Flash is extremely competitive.
giancarlostoro 2 hours ago||
Call me crazy but:

VRAM & Memory Requirements by Precision

• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).

• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).

• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)

VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.

Even so why would anyone not sleep on a model they cannot run?

kristopolous 2 hours ago||
Seriously, if a single politician stepped forward and said "i'll bring down ram prices" they could then shoot a puppy and call me a slur and I'd still go out and doorknock for them.

Memory companies have price fixed multiple times. They've paid hundreds of millions in fines. wikipedia even has a page on it. https://en.wikipedia.org/wiki/DRAM_industry_price_fixing.

Look at the financials of these companies, they're all making obscene margins and do they plan to increase production? No. Micron is doing a stock buy back to pump the price of their share.

The Micron CEO just recently said this is the exact plan https://www.theregister.com/systems/2026/10/01/ram-supply-se...

There's sanctions, tarrifs, and a DOJ who doesn't give a shit. Until we can fix that the insanity will continue. Phones will be unaffordable. Laptops will be obscene. Gaming consoles will be thousands of dollars. Desktops will be dead.

If you're waiting for some David Ricardo equation to happen, tough cookies, it's not coming.

The market is legally locked down and we're in hostage pricing mode.

And what's the story? You can't afford electronics because we're using it to build robots to take your job? I mean ...

Nobody is coming to save us. That's our job.

phil21 1 hour ago|||
> do they plan to increase production? No.

Micron has 3 brand new fabs currently under construction, 2 Boise, 1 in New York as the first of 4 planned for a campus.

Plus expanding other existing facilities.

These things take ~3-5 years from breaking ground to full production. You'd have had to anticipate the current demand years before it happened in order to be bringing production on-line before 2030 or so.

Samsung and HK Hynix also have fabs under construction and planned.

CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.

Not much you can really do to wish for more fabrication to exist on any timeline not measured in fractional decades.

Could they do more and react quicker? Probably, but everything I've read on the subject seems to point to 3 years is absolute bare minimum if you happen to have a shovel ready project with the land bought, local permitting completed, infrastructure extended to the site, and a skilled workforce already in place. They could suspend buy-backs/dividends today and dump it all into building production and there would be no material impact until around 2030.

> The Micron CEO just recently said this is the exact plan

CEO simply stated the demand pressure will not go away through 2027, and supply will not increase until around 2028 when currently under construction fabs start shipping volume. The article does not support your statement.

ttul 1 hour ago|||
Stanford tracks RAM prices in this nice little site: https://dam.stanford.edu/memory-prices.html

Costs did go nuts, but there are signs of easing in the market of late. CXMT is starting to have an impact and priced will probably fall in 2027.

kristopolous 1 hour ago||
https://pcpartpicker.com/trends/price/memory/ is better. pcpartpicker has the data.
cogman10 57 minutes ago||||
The second Micron boise fab hasn't even broken ground yet, they are still working on the first one. So don't expect these things to be completed in parallel.

Some of my family is pretty happy, though, with the job security as they are pretty convinced these projects are all going to take much longer than what's being stated publicly. Micron is saying the first chip from the new fab will be in 2027... though they also predicted it'd be 2026. The date seems pretty slippy.

dboreham 51 minutes ago||
Anyone who has been around the semiconductor industry since the last century will remember various huge fabs e.g. in Arizona that were partially built but never finished due to oversupply by the time the walls and roof were done.
BizarroLand 1 hour ago|||
Yeah, but why would they make consumer memory when HBM for GPUs is much more profitable?
kristopolous 55 minutes ago||
Capitalism eats itself this way. Second and third order effects will collapse the demand.

You need to keep the market healthy, not some insane Bitcoin style HODL pump - that's how you get wrecked.

I mean I'm not a neoclassicalist but I've read all of them. I'm in consensus with them here. There's a bunch of theories on what a healthy market is but what we're currently seeing matches none of them.

It's short term profitable but long term disastrous, especially in a world where new mathematics and techniques could literally collapse the demand overnight.

Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!

Some clever trick about how attention heads and context Windows work could potentially slash a bunch of requirements by giant margins and all they're doing is firing the starting gun at that global race with every obscenely priced unit they sell.

But if prices were reasonable, this wouldn't be an apocalypse. It'd be fine. Consumers wouldn't rush to 64GB, they'd say " Cool I can multitask now at 256" or " great I can do horizontal scalability' or something else.

But no they created the market conditions so now what would happen is the consumer will immediately flip the 192GB they don't need on eBay, hoping to snatch a profit before the prices tank and the second hand market will be flooded the rug will be pulled out from the luxury pricing and everyone will get screwed.

This has happened in electronics markets before. Many times.

When Engels talked about the grave diggers of capitalism they were looking at it through a 19th century labor/manufacturing lens but arguably this same dynamic is at play here.

usefulcat 1 minute ago|||
> Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!

If you were a DRAM manufacturer, isn't this exactly the kind of thing that would make you think twice about investing years and $billions in new fab construction?

Analemma_ 50 minutes ago|||
What "second- and third-order effects" do you suppose will collapse the demand for RAM? The people complaining most loudly about RAM costs are the people who want to run local models; if that becomes popular it will supercharge RAM demand, because locally-hosted models can't parallelize runs from many users the way cloud-hosted ones can. I don't see any slackening in RAM demand at any point in the foreseeable future, even if the big AI companies all go bust.
kristopolous 31 minutes ago|||
This is all hypothetical and debating hypotheticals isn't productive so let's roll back to markets.

Let's say ram used to cost $100 and now that same unit costs $1000. You paid say $500x1,000 for that unit during the price increase or some price where you can currently flip for profit.

You have a very expensive data center and you're in debt financed on the premise that you have these special computers.

Now a new technique comes out and it turns out you only need 1 memory unit for something that used to require 8 or 4 or some meaningful multiplier.

This stuff happens all the time. It's why we don't use BMP files on websites or serve giant MOV files on YouTube. It's why postgres queries are faster now than they were 10 and 20 years ago.

You rent out your machines. You need to service your debt.. Demand may 8x overnight to accommodate but you have a monthly bill to pay and that's unlikely. It's likely going to drop.

Think about it. Your customers are paying maybe $10,000 a month and serving their customers. Now they can drop that to $1,250.

On market if you were to sell some of that ram you have 100% profit right now but not for long.

Jevons paradox assumes unlimited capitalization, zero debt servicing, infinite time horizons...

We live in the real world so what do you do?

Historically the answer has been "sell that shit"

There's an aphorism for this "stairs on the way up elevator on the way down"

If we had a healthy market with sane prices where you can't flip the thing you bought for 100% profit the answer would be "create more value."

lxgr 25 minutes ago|||
> locally-hosted models can't parallelize runs from many users the way cloud-hosted ones can

Why not? Unlike many other workloads, LLM inference actually seems pretty suitable for decentralization (effectively stateless means no availability concerns; bandwidth and latency are relatively forgiving too).

Analemma_ 13 minutes ago||
I think locally-hosted models at the org level will definitely be somewhat popular, but you seem to be talking about decentralizing for people's personal, non-business use, and I just don't think that's going to happen to any real degree.

People who say they want local runs really mean it: they want local runs on hardware in their room, not on some decentralized system which, if it existed, would almost certainly just be a worse, less-reliable version of cloud hosting. I'm not saying nobody would use it, but it sounds a lot like things like IPFS, which have also completely failed to displace either cloud storage or buying a bunch of disks for your own private use.

ashdksnndck 1 hour ago||||
RAM manufacturers are bidding against NVIDIA and everyone else for the same constrained supply of EUV machines. And it takes years to build more fabs. Micron has multiple fabs coming online in 2027 and 2028.
m463 1 hour ago||||
> "i'll bring down ram prices"

wonder what voting would be like?

gamer vote ++

datacenter hater vote --

datacenter lobby ++

micron lobby --

rezonant 48 minutes ago|||
Yep, that's all the voting blocs.
mwambua 1 hour ago|||
Wouldn’t cheaper memory make it easier to bring compute out of data centers and onto consumer hardware?
TeMPOraL 1 hour ago||
Datacenter haters will read this as "that's still evil AI", and everyone else hopefully can count and understands it'll be worse for environment.
neya 55 minutes ago||||
> they could then shoot a puppy and call me a slur

I know it's just a figure of speech, but damn. I laughed out aloud in public just reading this.

bob1029 41 minutes ago||||
If we take some time to understand how HBM memory is manufactured (with particular focus on yield risk for final packaging steps), we will hopefully learn that the current capacity crisis is not bullshit.

I guarantee Micron & friends are not intentionally orchestrating their business such that they would suffer a massively reduced chance of yielding on a per-die basis. Unless someone is actually buying HBM devices, they are not going to be making them. These are not a commodity that can be speculatively manufactured in any economically rational way.

xyzsparetimexyz 1 hour ago||||
Neither political party cares at all about memory pieces get real lol
kristopolous 1 hour ago||
wait until holiday shopping...it affects the price of almost everything with a battery or power cord.
xyzsparetimexyz 10 minutes ago|||
What do you think they'll do? Neither repubs nor dems will touch ai companies in a meaningful way. Anything China does wrt memory fabs week be more significant
gchamonlive 1 hour ago|||
[flagged]
Analemma_ 1 hour ago|||
I don't think RAM vendors have formed a cartel and I think this is knee-jerk anger without any thought. RAM is a commodity product with massive upfront capex costs, and those always have boom-and-bust cycles. At various points in the 2010s and 2020s RAM vendors were getting eaten alive by a supply glut, this would not have happened if they were a cartel.

Is it really so hard to believe that RAM prices are up because demand is simply exceeding supply, especially in a market where additional supply takes years and billions of dollars to come online? There's no need to posit cartel behavior and a fair amount of evidence that there is none.

boustrophedon 5 minutes ago|||
The RAM vendors have formed cartels previously and been convicted, so although demand is exceeding supply it is not that crazy to at least consider.
kristopolous 1 hour ago|||
The AI boom started in 2022. Prices rose THREE years later after 2025 sanction and tariff style legislation to protect the market during a price hike.

I got a 4090 in 2023 for 1600, a 5090 in 2025 for 2000 with 256 DDR5 for about $1,000 ... and then, after some protectionist legislation passed, these prices quickly shot to the moon.

Connect the dots.

Analemma_ 44 minutes ago||
Man I think you're just spewing word salad and a lot of what you've written is either wrong or not even wrong. The AI hype really got started in 2022, but hype on social media doesn't mean anything for RAM prices, only real buildouts do that. They rose pretty steadily until OpenAI revealed their shenanigans re: locking up a ton of supply from two different vendors with secret contracts, and that's when the takeoff really started. This is definitely scummy behavior from OpenAI (big surprise), and I actually think they arguably should see an antitrust investigation for that (not that that will ever happen), but OpenAI is a buyer; that's not the same thing as the vendors forming a cartel.

You can't say "connect the dots" at the end of a raving, mostly-incorrect post and act like you've made an ironclad argument.

petu 1 hour ago|||
There's no BF16, original full quality weights are quantized already and 510GB.

Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.

Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).

mrinterweb 37 minutes ago|||
Projects like DwarfStar https://github.com/antirez/ds4 really lower the hardware bar a lot so Deepseek 4.1 flash and other mixture of expert models can run on consumer hardware. There are also other inference providers who make their money serving openweight models. Services like OpenRouter make it all too easy to utilize these models. Access to these models isn't hard. The hardware moat is becoming pretty easy to bridge.
contingencies 9 minutes ago||
More concretely DwarfStar M5 128GB Deepseek 4.1 flash 1K tokens @ 29s, 5K tokens + reasoning @ 147s, 10k token prompt @ 463 tokens/s = 22s. Hardware buy-in USD$7K / AUD$8.5K / EUR€6.8K. At typical workloads, ROI is still poor vs. current-era subsidies, but owning hardware is good for privacy/longevity/connectivity independence. Whether you actually consider Apple hardware 'owned' is a valid and thought provoking question.
onlyrealcuzzo 3 minutes ago||
Still gonna take 2-3 years to get DeepSeek V4.1 Flash quality at decent speeds on reasonably priced hardware.

Hardware update cycles are 2-3 years even on the high end, so it's still a ways away before "good enough" and "local" belong in the same sentence for the average person.

And by then, DeepSeek V6 Flash will be too cheap to meter, 5x faster, and 10x better, so... You'd still need to go out of your way.

Most people are spending most of their time on their phones anyway. ..

wren6991 22 minutes ago|||
Are you counting the n-gram/PLE as part of the model weights there? They can go in host memory. Would be good to show your working. Also the released weights are pre-quantised and presumably QATed, so your "Full Precision" and INT8 are simply not a version of the model that actually exists.

Edit: I went and checked for you. The LM backbone is 307.2 GB (286.1 GiB), straight from DeepSeek's upload. The n-gram table is 203.1 GB (189.1 GiB), which goes in host RAM. Note the embeddings are higher precision than the expert tensors, so it's a larger fraction of the bytes than it is of the parameters.

So,

> Call me crazy but:

You're crazy. :-)

ct520 30 minutes ago|||
"1070s or 1070 TIs because GPUs have been severely overpriced for too long" ... ."

1070ti launch MSRP was $450 ish. 5070 could be had in the last year for 5xx-6xx range easily.

All things considered - (inflation being about 30%~ (guess)) between these two timelines. You are looking at 300% performance difference at a cost dollar for dollar that is cheaper then when they purchased their cards.

Might be a bit of a stretch blaming it on "severely overpriced for too long..."

ManuelKiessling 51 minutes ago|||
Thanks for the data!

Allow a question from someone who’s only got a very vague idea of how this kind of stuff works behind the scenes: say I rent usage of this model through one of the many LLM hosting providers out there, and let‘s assume I use it extensively through something like Pi or OpenCode and vibe code away all the time, keeping the hosted model occupied as much as I can, happily burning my credits.

Does that mean that there is a hardware cluster as described by you above that is crunching away just for me?

So at FP16, I alone keep a 1,664 GiB system occupied all the time?

rnxrx 9 minutes ago|||
It depends hugely on what "rent usage of this model through one of the many LLM hosting providers" means. If you're asking them to host the model privately then yes, all of that 1.6T of RAM is likely in use holding weights, activations and KV cache by an inference engine that's only getting/answering requests from you alone. When you aren't actively using the model the hosting process is still active and waiting with all of that memory still wired to it.

As background: For the most part VRAM oversubscription/paging/swapping isn't a thing in the same way that RAM for a VM often is. There are some approaches to it, but (to my knowledge) not at that sort of scale.

There are some systemic reasons for this, but very broadly speaking the GPU vendors are building toward the highest bandwidth and lowest latency possible, and the overhead/complexity of something like protected memory modes serves neither of those priorities.

DrammBA 42 minutes ago|||
No, a cluster can server multiple users at the same time, providers cap the tok/s so that one cluster can run inference on multiple inputs at the same time. OpenAI with their new ultrafast mode is probably reserving the whole cluster or prioritizing requests of ultrafast users above others with a higher tok/s hence the high price and high speed. There's many other knobs providers tweak that they don't show the users, for example I doubt many providers are hosting the full FP16 version.
keammo1 1 hour ago|||
The article isn't just about running locally though. The author is saying it's super cheap to run the model through Opencode Go (and presumably OpenRouter etc.) Personally I'm always most excited by models I can actually run locally, but even these huge open source models open up the competitive landscape for companies to let you call models via an API or just lease compute. And they don't have to charge you to offset research, training, huge staffs of the best minds in the world, crazy PR etc. I think that's a big win for customers and buts competitive pressure on the frontier labs as well.
crossroadsguy 1 hour ago|||
I did somet math and completely gave up on the idea of trying any worthwhile local model and figured I'd rather pay the 15-30 USD per month via subscription and/or API key combos for years than buying a local setup which might go out of date very fast, if it doesn't goes kaput just out of warranty. I won't be surprised if RAM scarcity is an concerted effort to herd people towards the remote models :)
girvo 1 hour ago|||
Not quite: not all of this needs to be in VRAM

It has a set of n-gram tables which you can stream from system RAM or even NVMe

That said it’s still quite big! I can’t fit it on my DGX Spark, though I believe you can if you have two?

jonsoft 27 minutes ago|||
It needs 3-4 Sparks to run well (at an acceptable quantization and sufficient KV cache):

https://github.com/christopherowen/spark-ds41f

girvo 6 minutes ago||
Ah that’s a shame. GLM 5.3 Flash is honestly as good IMO and can run on two pretty successfully from what I understand.

I’m quite spoiled with how good Qwen 3.8 Flash Next is on a single spark though: shocking how good local models are getting on attainable-ish hardware

rsolva 59 minutes ago|||
I have access to two and will explore this the coming weeks.
girvo 5 minutes ago||
Also give GLM 5.3 Flash a try: it’s shockingly good too in my testing, and I believe eugr has a TP=2 recipe to use for sparkrun
ByteAtATime 1 hour ago|||
I think, considering the size of this model, it's closer to a Pro than a Flash on everything other than speed
apitman 1 hour ago|||
> Even so why would anyone not sleep on a model they cannot run?

Because it's an open model so providers compete on price.

jauer 44 minutes ago|||
This “blame sama for memory prices” meme is so tired.

He gave demand signal so many times years ago and was mocked for it and now we have the consequences of industry not taking him seriously.

anvuong 1 hour ago|||
I just un-retire my pair of 1080Ti for some small models development because the current GPU prices literally make me sad.
cookiengineer 1 hour ago|||
It's dangerous to go alone. Take this: [1]

I reimplemented most of the features of the Deepseek v4.1 flash paper (apart from quantization aware training which doesn't make sense because my implementation uses float32 precision anyways)

I'm currently learning how to distill reasoning traces (check my other github repositories) but I think that a locally selfhostable deepseek is possible with my mixture of experts sharding mechanism. I decided to optimize everything for CPU parallelization, with the idea that the KV cache and meta model have to run from CPU RAM anyways, so the experts can also be loaded/unloaded at runtime if needbe, to save more RAM.

My assumption is that the KV cache optimizations in combination with the CED and compressed attention features are the reason why v4.1 flash has so few hallucination problems and such a strong self-lookup/thinking behavior. But that's more a gut feeling, need to evaluate and test this more thoroughly.

Anyways, would love to see someone train this on their own datasets. Currently my pipeline is kinda optimized for parquet and zim files.

[1] https://github.com/cookiengineer/gonano

CamperBob2 1 hour ago|||
You can run it locally for the price of a decent car, or run it (hopefully) privately on somebody else's hardware at vast.ai or a similar provider for much less. What's not to like?

No, you won't get frontier-level intelligence on a 1070Ti. Yes, it should be illegal to do what Altman did. Since we clearly don't live in the best of all possible worlds, we need to settle, and DS4.1 Flash is a good place to do that.

For tasks that don't require vision I personally like the NVFP4 quant of GLM 5.3 from Local Inference Lab better than DS4.1F, but they are both well beyond awesome.

nullc 1 hour ago|||
The bulk of its weights are natively MXFP4. And engram values don't need to be in vram.
functionmouse 1 hour ago|||
one can make a fine gaming pc for ~$350

1660 ti, 4790k, 16gb ddr3

holoduke 1 hour ago|||
He doesn't ruin the cost of memory. Advances in memory size and speed are now in full speed mode. Expect drastic increase in the upcoming years. Big factories are in the making and planned. Gigalab in the US and many others in the east. Since 2010 we have computers with 16gb as being normal. Finally we are moving into a new era where the standard will be 64gb next year and 128 in 2028. Hopefully we reach 1tb in 2030.
CorrectHorseBat 1 hour ago||
I've read the exact opposite, vendors are reducing the standard from 16GB back to 8GB
holoduke 1 hour ago||
That's only temporary till production meets demand again.
CorrectHorseBat 56 minutes ago||
Which is not going to happen in the upcoming years
liuliu 1 hour ago||
What are you talking about? The model is native NVFP4, why you run it at any precision higher than that?
ne01 39 seconds ago||
Deepseek V4.1 Flash is a hidden gem, really. Not to mention, you can easily get it through many providers that offer zero data retention and consistent speeds above 200 tokens per second!
mlinsey 1 hour ago||
I'm paying for the heavily-discounted subscriptions, not the API rates. There isn't really a cost gap for me. DeepSeek doesn't have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100/mo Z.ai subscription compared to Opus 5.5 on a $100/mo Claude subscription, there wasn't a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot closer, but I didn't use it enough to really say for my workloads).

I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).

Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.

apitman 55 minutes ago|
The tightening of subscription value has already begun. dsv4.1f is already worth paying for at market API prices. Maybe it goes to 2x because apparently no one has figured out how to match DeepSeek's insane caching efficiency, but I don't see it getting much worse than that.

Plus you can also get dsv4.1f subsidized. OpenCode Go gives 4x if I understand their pricing correctly. Anecdotally, I feel like I get way more out of my $10/mo OpenCode Go sub for the price than my $20/mo ChatGPT, even using gpt-6.1-sol high which is very cheap, and I have yet to convince myself dsv4.1f is a worse model.

rapind 42 minutes ago||
It's really not cheaper than frontier subscriptions. It's getting closer, and it's a great model, but it is not more value per task than the frontier subscriptions. Don't be swayed by the token costs, it's very chatty, like 3x more tokens for the same task as sol. I used dsf 4.1 full time for about a week.

It blows frontier API pricing out of the water, but again, look at cost per task, not token usage. Still easily wins though for my work.

I do think it's the most viable alternative I've seen so far, and that applies pressure to the frontier models. Should subscription prices hike or become unavailable for some reason, I know what I'll be using.

When pricing this, it's important to consider whether or not you want to opt out of data training. You won't get the advertised rate. Also the dsf 4.1 subscription providers are throttled af... and of course they are, because otherwise they'd be haemorrhaging money.

apitman 39 minutes ago|||
I am looking at cost per task (and speed per task), both benchmarks and anecdotal experience.
LeBit 5 minutes ago|||
What are you talking about?

DS4.1 Flash not really cheaper than frontier models???

It is insanely cheaper.

gregwebs 15 minutes ago||
I have been using DeepSeek 4.1 flash intensively for over a month. If I run it all day long it costs $1-2. Its fast. Previously I was always quickly running up to my Claude/Codex 5 hour window (on the $20/month plan). The cost savings of DeepSeek is real as shown in this article and I am using subsidized plans.

DeepSeek is horrible at grilling sessions (the /grill* skills to make technical decisions). It doesn't know how to explain things. Maybe the skill could be adjusted. It also doesn't come up with as good solutions as Opus/Sol.

What I use it for is

  * the orchestator of my coding workflows
  * the tester/verifier of code changes
  * the sub agent that explores code or does web searches
  * putting together code base research reports
Previously I planned with Opus/Sol/Astra and then I used DeepSeek for coding, and then reviewed with Opus/Sol/Astra. With the cost improvements to Opus/Sol I am trying to use them for coding instead now so there will be less back and forth review needed.

They are all working together in Pi using the extension @tintinweb/pi-subagents where my workflow skill is calling different subagents that use different models.

Luna is cost competitive, but doesn't score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn't seem like a competitive price/performance mix.

oh_no 5 minutes ago|
Luna is 1 point being on AA's index at 1/4 the cost, yes it "doesn't score as well" but paying 4x for 1 point is crazy if you're going off benchmarks.

AA has Haiku 5.5 as cheaper than 4.1 Flash (both on Max, which isn't ideal but what can ya do) and a 4 point intelligence gap.

Why do people like to think open models are more competitive than they are?

p1necone 1 hour ago||
I have a pretty large, complex project I've been building with heavy AI use (new language + compiler). I was following a 'strong model as orchestrator launching cheap models as implementers' pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).

I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.

However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.

rspeele 1 hour ago|
On the Claude side of things I was previously following "strong model directs weak" with Fable directing Opus/Sonnet (its choice per-task). Since Opus 5.5 came out I've just been having Opus direct Opus.

The sub-agent separation is still valuable to keep context clean for the orchestrator, but I just have no reason to use Sonnet as the grunt-work implementer because I'm finding it hard to run out of tokens with Opus 5.5 on a $200 subscription plan. It's really really good at subjective quality of work per token used.

p1necone 1 hour ago||
I would probably go that route if I could use other harnesses with claude models, but I don't want to be locked in to claude code, and their API pricing (which you need to use it with other harnesses) is so much higher than subscription.
lmf4lol 1 hour ago||
Oh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.

Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.

But as a main driver. I love flash. And it brought our bill down by A LOT :D

crossroadsguy 56 minutes ago||
What is the cost of access like for DeepSeek-v4.1-flash, compared to GLM-5.3-flash via ZAI's Coding Plan? Because that's what I use; and often hit the "wait". I wouldn't mind trying a new model subscription or even API access which hits around glm-5.3-flash level weight class (which seem to be enough for me; with quite some human suprvision and nudging) but gives muuuuuuch moooore tokens for the same price.

(And what are the preferred providers?)

PcChip 1 hour ago|||
>We run all our Personal Assistants now on flash

are you worried about sending all your data to third parties, especially if they're in different countries?

techmunky 48 minutes ago|||
Not shilling for them but Ollama cloud hosts domestically with ZDR afaik. I run 95% of my open weight inference through them. The rest goes through Opencode Go $10 plan (which is enough to run 3 hermes agents on DSF 4.1 and leave plenty of left to experiment with when new models drop).
octoberfranklin 43 minutes ago||
Just a reminder that any API using a Cloudflare TLS certificate isn't ZDR.

The model engine provider might be ZDR, but the service as a whole isn't.

techmunky 14 minutes ago||
did not know. ty! have my updoot as thanks
figmert 29 minutes ago||||
I use it through OpenRouter, which has ZDR enforcement.
crossroadsguy 55 minutes ago||||
I have asked OP that question but I think there are providers who are not in China and they just host the model/inference.
yieldcrv 55 minutes ago|||
Just use a provider hosting it in your country especially if your country has major data centers then its the same as using Anthropic or GPT of GCP Model Garden or AWS Bedrock

nobody here is talking about running frontier level intelligence locally so if you’re Chinaphobic and prefer layers of corporations siphoning your data in between you and the party there are plenty of options instead of directly to the party

aftbit 1 hour ago||
Have you compared it against actual SOTA models like latest Fable or Astra?
mtrovo 23 minutes ago|||
The author explains this very well tbh:

> Today's models are now good enough for high-quality unattended tasks. Chasing the latest and greatest is silly. It is fun to see the new Fable capabilities, but the tasks we throw at them are usually ridiculous (maybe even insulting) if you believe in LLM sentience. It's like asking a math PhD to organize the files on your desktop.

I'm using DS V4.1 Flash as my main model since their release and it works great for all my coding tasks. My setup is OpenCode Go subscription and obra/superpowers skill.

The only times I try to change models are on general planning tasks (like research this codebase for tech debt mitigation opportunities) or if I need deep research which would benefit from searching the web, in which I still think Gemini is still the best because of the speed and access to google search index. But these are not even 20% of my daily tasks.

sneurlax 1 hour ago|||
Of course there's still a huge performance gap

but DS 4.1 Flash is good enough for most tasks

james2doyle 14 minutes ago||
Been using Flash 4.1 via the ante harness to blast through a GBA recomp. The ante team has pushed hard to make Flash 4.1 perform well under it. So far, I've maybe spent $10 over the last 3 days. Its a real workhorse and works much better in this harness
_jayhack_ 4 minutes ago||
Enterprise is not freaking out because DeepSeek 4.1 Flash does not actually occupy a spot on the Pareto frontier for non-coding enterprise workflows. We see this at my employer, focused on non-technical knowledge work. Luna 6 and now Haiku 5.5 are both very competitive if not better on all axes that we care about
wren6991 7 minutes ago|
It's a solid little model, and I appreciate DeepSeek's commitment to the bit in releasing a brand new pretrain, double the size, numerous architectural innovations as a ".1" release over the excellent DeepSeek V4 Flash.
More comments...