Top
Best
New

Posted by jonotime 2 days ago

Why isn't the industry freaking out about DeepSeek 4.1 Flash?(www.dgt.is)
1090 points | 959 commentspage 3
apitman 1 day ago|
> With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited

My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).

I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.

This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.

youniverse 1 day ago||
How are you guys doing orchestration? I have been fumbling around in my free time trying to build something for myself but is there a repo or something that just works?
apitman 1 day ago|||
I might not be the best person to ask. I use Pi harness in tmux and just ask my current agent to spawn interactive pi instances in new tmux windows, create a sentinel file for each of them, and monitor the sentinel files for signals every 2 seconds.

Currently have auto compaction turned off. When the orchestrator's context is getting close to full, I have it write a handoff markdown file and point a fresh agent at it.

I do feel like I'm getting close to the point where I might be ready for something more sophisticated, especially wrt to subagents communicating with the orchestrator.

crossroadsguy 1 day ago|||
That limitation is what stops me from using Pi for anything serious. I may have to configure it and then configure it and it will eventually become a codex, a claude code or so. I recently heard the maintainers added mcp to it (in stock, not via plugin), I wonder what stopped them from adding subagent function, and decent loop capacity to it.
mappu 22 hours ago|||
> Currently have auto compaction turned off. When the orchestrator's context is getting close to full, I have it write a handoff markdown file and point a fresh agent at it.

In a good harness that should be how auto compaction works anyway

apitman 13 hours ago||
Yes, and I'll probably trust it eventually. But currently I always want to know when compaction is happening, so I can correlate any drops in quality or weird behavior.
CharlieDigital 1 day ago|||
Many orchestrator harnesses exist.

Check: https://agentmgmt.dev/ and find the one that works for you.

I quite like Paseo (been maining it for a week), but Orca also looks good.

sourcecodeplz 23 hours ago||
you can spend your opencode go quota faster than 1 month. with the limits i think its about two weeks.

so even if the week reset with some %usage left, its not actually lost if its not the end of the month.

skeptic_ai 18 hours ago||
I reverse engineer a basic flash game and spent 3 open code accounts.
alex-moon 1 day ago||
I think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.
rhdunn 21 hours ago|
There are also several issues at play here:

1. a model that works for one person/task may not work for another;

2. there are many models (DeepSeek, Qwen, GPT, Claude, Gemini, etc.) that are released every 6 months or so;

3. it takes time to use, test, and evaluate the suitability of a new model and not everyone has an automated evaluation process for their use cases.

Thus, if you find a model that works for you then you are not going to spend more time evaluating a model that may not work, or may only do so when time permits.

leptons 18 hours ago||
I'm pretty happy with claude $20/mo for personal, and $200/mo at work. It does everything I need it to, at a price that I can aaccept. If they raise the price, then I'll start looking around at the other options. I do not like Altman or Musk, and I don't want to get involved with China, so I'd rather avoid those products. But I do know that prices will have to increase at some point, so I'm still hand-coding a few projects to keep my skill level up should LLMs just not be affordable in the future.
K0IN 10 hours ago||
I really really love deepseek v4(and 4.1), but for everything I thrown at it, it felt like gpt 5.6 + 6 luna or terra do it faster, in way less tokens and in a way I like more. So even tho the price is (very) cheap, i found myself using it less just cause I don't want to wait on something I have to itteratee on with the model.
swiftcoder 1 day ago||
I think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 even cheaper so long as you are willing to share data with them
9dev 1 day ago||
Whatever the question, Meta is the wrong answer.
jacquesm 1 day ago||
> so long as you are willing to share data with them

I don't think so.

010ED67913 20 hours ago||
yeah this is hacker news. we only share our data with Dario, Altman, Musk, and the CCP. Not untrustworthy people like zuckerberg.
jacquesm 20 hours ago||
I don't share my data with any of those either.
swiftcoder 12 hours ago||
That leaves you with who, exactly, in the LLM game? Google? Mistral?
weknowbetter 10 hours ago||
People who have not used DeepSeek massively under estimate it's capability and how little it costs to run.
saberience 10 hours ago|
Yes but they're not under-estimating how much worse it is than Opus 5.5 or Astra...
Zambyte 9 hours ago||
But they're almost certainly overestimating the level of intelligence required to comfortably do their tasks.
cesarvarela 6 hours ago||
The issue is time. Even if DeepSeek matches Opus on 9 out of 10 tasks, failing on the tenth can cost you more in lost time than you saved in money.
kraig911 9 hours ago||
Everyone I talk to about Deepseek has a bad taste it feels like for CHinese models? I have a Kimi sub and use deepseek a lot at home. I've hit my limit on K3 many a time for Deepseek to come in and finish a job. So far no complaints. I feel it makes a lot of round trips by design but over all a good option. It's hard to compete though with an open sub out on codex/claude and just do everything in that subscription. Deepseek for me is when I run out of my subs. It's usually always coming behind a great session and fixing things so I don't give it a chance to try something novel on it's own.
zug_zug 1 day ago||
I did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.

That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.

0xbadcafebee 1 day ago|
I keep a spreadsheet that estimates actual value (dollar amount per token per month, per subscription rate limit) and open weights are basically always cheaper than frontier weights. Recently things like GPT 5.6 Luna finally got the frontier close to the value of open weights but their limits keep them behind.
Primer81 16 hours ago||
I've been waiting for something like this article for a month. The other AI companies are being left in the dust. I rip about 200 million tokens at least a day and spend maybe $3 with deepseek flash v4.1 for self hosted related coding tasks for ~20 projects i work on / maintain in parallel. Its a daily driver for sure. not even worth considering anything else, but maybe a locally run model on my GPU at the moment.
quietfox 15 hours ago|
What’s your dev environment with flash v4.1? Do you call the API directly or via Openrouter or something else?
giancarlostoro 1 day ago|
Call me crazy but:

VRAM & Memory Requirements by Precision

• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).

• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).

• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)

VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.

Even so why would anyone not sleep on a model they cannot run?

kristopolous 1 day ago||
Seriously, if a single politician stepped forward and said "i'll bring down ram prices" they could then shoot a puppy and call me a slur and I'd still go out and doorknock for them.

Memory companies have price fixed multiple times. They've paid hundreds of millions in fines. wikipedia even has a page on it. https://en.wikipedia.org/wiki/DRAM_industry_price_fixing.

Look at the financials of these companies, they're all making obscene margins and do they plan to increase production? No. Micron is doing a stock buy back to pump the price of their share.

The Micron CEO just recently said this is the exact plan https://www.theregister.com/systems/2026/10/01/ram-supply-se...

There's sanctions, tarrifs, and a DOJ who doesn't give a shit. Until we can fix that the insanity will continue. Phones will be unaffordable. Laptops will be obscene. Gaming consoles will be thousands of dollars. Desktops will be dead.

If you're waiting for some David Ricardo equation to happen, tough cookies, it's not coming.

The market is legally locked down and we're in hostage pricing mode.

And what's the story? You can't afford electronics because we're using it to build robots to take your job? I mean ...

Nobody is coming to save us. That's our job.

phil21 1 day ago|||
> do they plan to increase production? No.

Micron has 3 brand new fabs currently under construction, 2 Boise, 1 in New York as the first of 4 planned for a campus.

Plus expanding other existing facilities.

These things take ~3-5 years from breaking ground to full production. You'd have had to anticipate the current demand years before it happened in order to be bringing production on-line before 2030 or so.

Samsung and HK Hynix also have fabs under construction and planned.

CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.

Not much you can really do to wish for more fabrication to exist on any timeline not measured in fractional decades.

Could they do more and react quicker? Probably, but everything I've read on the subject seems to point to 3 years is absolute bare minimum if you happen to have a shovel ready project with the land bought, local permitting completed, infrastructure extended to the site, and a skilled workforce already in place. They could suspend buy-backs/dividends today and dump it all into building production and there would be no material impact until around 2030.

> The Micron CEO just recently said this is the exact plan

CEO simply stated the demand pressure will not go away through 2027, and supply will not increase until around 2028 when currently under construction fabs start shipping volume. The article does not support your statement.

ttul 1 day ago|||
Stanford tracks RAM prices in this nice little site: https://dam.stanford.edu/memory-prices.html

Costs did go nuts, but there are signs of easing in the market of late. CXMT is starting to have an impact and priced will probably fall in 2027.

kristopolous 1 day ago|||
https://pcpartpicker.com/trends/price/memory/ is better. pcpartpicker has the data.
fabioborellini 18 hours ago|||
You can see the SKUs they base the prices on by hovering over a data point, and the products just don't generally represent the market. Oftentimes the whole market is represented by some bottom of the barrel legacy memory, single stick in some warehouse. During the last years DDR3 and DDR4 haven't gotten as expensive as the current generations, and I think the memory compatible with current systems should represent the market rather than a Toshiba Satellite 4 GB extension kit.
GeekyBear 1 day ago||||
> CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.

It's taken them this long to catch up to the DDR5 standard. They've only recently been through qualifications to be a DDR5 supplier for the big boys.

> Every Major Motherboard Maker Now Validates CXMT DDR5

https://www.techtimes.com/articles/321572/20260725/every-maj...

After their recent IPO, they have more than enough cash to ramp up in a major way.

It's just a matter of time.

kristopolous 6 hours ago||
I'm fully open to the idea of having to fly to China to buy hardware. If we're seeing $25,000 US prices versus like $4,000 in China, might as well go visit the great wall and have a vacation.
cogman10 1 day ago||||
The second Micron boise fab hasn't even broken ground yet, they are still working on the first one. So don't expect these things to be completed in parallel.

Some of my family is pretty happy, though, with the job security as they are pretty convinced these projects are all going to take much longer than what's being stated publicly. Micron is saying the first chip from the new fab will be in 2027... though they also predicted it'd be 2026. The date seems pretty slippy.

minraws 20 hours ago|||
If memory prices cool, in about 2 years, Micron will stop new projects, they have done it before.

Especially given CXMT has been able to scale up much faster than what most people expected, only reason their isn't a bigger impact is modern HBM is hard to CXMT even today.

We are likely to see supply double in the next 3 years, but demand even out with optimizations, cooling of data center demand, and most importantly moving some of the dram to flash demand instead which is much easier to produce and scale.

dboreham 1 day ago|||
Anyone who has been around the semiconductor industry since the last century will remember various huge fabs e.g. in Arizona that were partially built but never finished due to oversupply by the time the walls and roof were done.
BizarroLand 1 day ago|||
Yeah, but why would they make consumer memory when HBM for GPUs is much more profitable?
kristopolous 1 day ago||
Capitalism eats itself this way. Second and third order effects will collapse the demand.

You need to keep the market healthy, not some insane Bitcoin style HODL pump - that's how you get wrecked.

I mean I'm not a neoclassicalist but I've read all of them. I'm in consensus with them here. There's a bunch of theories on what a healthy market is but what we're currently seeing matches none of them.

It's short term profitable but long term disastrous, especially in a world where new mathematics and techniques could literally collapse the demand overnight.

Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!

Some clever trick about how attention heads and context Windows work could potentially slash a bunch of requirements by giant margins and all they're doing is firing the starting gun at that global race with every obscenely priced unit they sell.

But if prices were reasonable, this wouldn't be an apocalypse. It'd be fine. Consumers wouldn't rush to 64GB, they'd say " Cool I can multitask now at 256" or " great I can do horizontal scalability' or something else.

But no they created the market conditions so now what would happen is the consumer will immediately flip the 192GB they don't need on eBay, hoping to snatch a profit before the prices tank and the second hand market will be flooded the rug will be pulled out from the luxury pricing and everyone will get screwed.

This has happened in electronics markets before. Many times.

When Engels talked about the grave diggers of capitalism they were looking at it through a 19th century labor/manufacturing lens but arguably this same dynamic is at play here.

Analemma_ 1 day ago|||
What "second- and third-order effects" do you suppose will collapse the demand for RAM? The people complaining most loudly about RAM costs are the people who want to run local models; if that becomes popular it will supercharge RAM demand, because locally-hosted models can't parallelize runs from many users the way cloud-hosted ones can. I don't see any slackening in RAM demand at any point in the foreseeable future, even if the big AI companies all go bust.
sieve 1 day ago|||
> The people complaining most loudly about RAM costs are the people who want to run local models

This is a tiny percentage of the population.

Samsung is cutting phone production because of RAM prices.[1] The consumer market is badly affected: budget phones, laptops, general electronics.

The budget segment of sub $100 devices in India has been almost wiped out. Manufacturers cannot afford to spend 50% BOM on RAM+storage. Unless employees are getting a 15-20% wage rise this year, I expect a similar situation in most places.

Between the engineered conflict in the ME triggering O&G price rises, and stratospheric RAM pricing, the situation is pretty bad.

[1] "There is no profit even if we sell"…Samsung to cut smartphone production by 30% (https://www.mt.co.kr/en/tech/2026/10/08/2026100709554237233)

kristopolous 1 day ago||||
This is all hypothetical and debating hypotheticals isn't productive so let's roll back to markets.

Let's say ram used to cost $100 and now that same unit costs $1000. You paid say $500x1,000 for that unit during the price increase or some price where you can currently flip for profit.

You have a very expensive data center and you're in debt financed on the premise that you have these special computers.

Now a new technique comes out and it turns out you only need 1 memory unit for something that used to require 8 or 4 or some meaningful multiplier.

This stuff happens all the time. It's why we don't use BMP files on websites or serve giant MOV files on YouTube. It's why postgres queries are faster now than they were 10 and 20 years ago.

You rent out your machines. You need to service your debt.. Demand may 8x overnight to accommodate but you have a monthly bill to pay and that's unlikely. It's likely going to drop.

Think about it. Your customers are paying maybe $10,000 a month and serving their customers. Now they can drop that to $1,250.

On market if you were to sell some of that ram you have 100% profit right now but not for long.

Jevons paradox assumes unlimited capitalization, zero debt servicing, infinite time horizons...

We live in the real world so what do you do?

Historically the answer has been "sell that shit"

There's an aphorism for this "stairs on the way up elevator on the way down"

If we had a healthy market with sane prices where you can't flip the thing you bought for 100% profit the answer would be "create more value."

charcircuit 1 day ago||
>Your customers are paying maybe $10,000 a month and serving their customers. Now they can drop that to $1,250.

Or they could stay at $10,000 per month since they are willing to pay that much already.m, so they just use AI more and in more places.

lxgr 1 day ago|||
> locally-hosted models can't parallelize runs from many users the way cloud-hosted ones can

Why not? Unlike many other workloads, LLM inference actually seems pretty suitable for decentralization (effectively stateless means no availability concerns; bandwidth and latency are relatively forgiving too).

Analemma_ 1 day ago||
I think locally-hosted models at the org level will definitely be somewhat popular, but you seem to be talking about decentralizing for people's personal, non-business use, and I just don't think that's going to happen to any real degree.

People who say they want local runs really mean it: they want local runs on hardware in their room, not on some decentralized system which, if it existed, would almost certainly just be a worse, less-reliable version of cloud hosting. I'm not saying nobody would use it, but it sounds a lot like things like IPFS, which have also completely failed to displace either cloud storage or buying a bunch of disks for your own private use.

lxgr 21 hours ago||
Some people will care a lot about keeping their data on-prem, but many others probably won't, and the former can then resell their spare capacity to the latter.

Decentralized storage is much harder, since there reliability matters a lot more as it's inherently stateful. You have to assume data loss, so you have to replicate everything; with inference, you only have to spend extra resources at failover time. Also storage can't be time-shared in the same way as compute; if it's full, it's full even when not actively accessed.

Analemma_ 10 hours ago||
I mean I guess there's nothing left which can settle this disagreement except to see how it turns out. I'll just restate my opinion that, from a user experience and product perspective, decentralized model hosting will look like using the big labs, except worse in every way. I don't expect it to please either the people who want local control or the people who are happy using the products from the big labs now.
usefulcat 1 day ago|||
> Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!

If you were a DRAM manufacturer, isn't this exactly the kind of thing that would make you think twice about investing years and $billions in new fab construction?

hn_acc1 1 day ago||||
>Seriously, if a single politician stepped forward and said "i'll bring down ram prices" they could then shoot a puppy and call me a slur and I'd still go out and doorknock for them.

How many people, outside of tech geeks and megacorps care about RAM prices? And how gullible would you be to BELIEVE the politician they could actually make it happen, and even if they did, that it would extend to the average person, and not JUST megacorps/megadonors?

Aerroon 1 day ago|||
Phone companies have been differentiating their models based on RAM for a decade. As have laptop and desktop sellers. The reason your router sometimes randomly crashes could very well be a result of not enough memory. The reason it takes such a long time to launch some programs repeatedly is because you don't have enough memory to cache it. Swapped from your browser to an app on your phone, but when you go back to the browser the site has reset and you lost everything you were working on? Not enough memory. Etc.

I think a lot of people care about the downstream effects of memory prices, but I agree with you that they may not realize that they happen because of memory prices.

gruez 1 day ago||
>The reason your router sometimes randomly crashes could very well be a result of not enough memory. The reason it takes such a long time to launch some programs repeatedly is because you don't have enough memory to cache it. Swapped from your browser to an app on your phone, but when you go back to the browser the site has reset and you lost everything you were working on? Not enough memory. Etc.

That might be true at micro level, but at the macro level more memory just means developers get more lazy with their optimizations, causing apps to get more bloated, eating up any gains in extra memory. There's no reason why slack needs 1+GB to run, yet people are perfectly happy to put up with it.

kergonath 15 hours ago||||
> How many people, outside of tech geeks and megacorps care about RAM prices?

They don’t care about RAM prices, but they do care about the price of things that have RAM in them (or even NAND), and all of them are increasing way faster than inflation.

Fricken 19 hours ago|||
Everybody has a phone, they're costing 25% more. GameStop is selling second hand PS5s for $1400. RAM prices affect most people.
ashdksnndck 1 day ago||||
RAM manufacturers are bidding against NVIDIA and everyone else for the same constrained supply of EUV machines. And it takes years to build more fabs. Micron has multiple fabs coming online in 2027 and 2028.
dboreham 1 day ago||
Don't believe Nvidia has any fabs of its own.
mcv 22 hours ago||||
I keep arguing that memory needs more competition, and people keep pointing out that it's too slow and expensive to ramp up. But if the threshold to enter that market is so steep, that means it cannot function as a free market and requires regulation.

In this case I think investment in more production is the only option, and it needs to happen even if it is expensive and slow.

m463 1 day ago||||
> "i'll bring down ram prices"

wonder what voting would be like?

gamer vote ++

datacenter hater vote --

datacenter lobby ++

micron lobby --

rezonant 1 day ago|||
Yep, that's all the voting blocs.
mwambua 1 day ago|||
Wouldn’t cheaper memory make it easier to bring compute out of data centers and onto consumer hardware?
TeMPOraL 1 day ago||
Datacenter haters will read this as "that's still evil AI", and everyone else hopefully can count and understands it'll be worse for environment.
absoluteunit1 18 hours ago||||
> Seriously, if a single politician stepped forward and said "i'll bring down ram prices" they could then shoot a puppy and call me a slur and I'd still go out and doorknock for them.

I spit out my coffee laughing when I read this

BatteryMountain 22 hours ago||||
The situation is actually much worse and the long term consequences will start materializing soon. The wholesale theft of humanities soul is in progress. It won't be a pretty sight in supposedly civil first world countries, when the human spirit awakens. Currently we are still pressing that snooze button hard and repeatedly, as I think most are keenly and deeply aware of what needs to happen but that too will cost our souls.
chrismsimpson 21 hours ago||||
> they could then shoot a puppy and call me a slur and I'd still go out and doorknock for them

Priorities

j16sdiz 19 hours ago||||
> Seriously, if a single politician stepped forward and said "i'll bring down ram prices"

How? Increase production? The time needed to scale up the production is longer than one election cycle.

neya 1 day ago||||
> they could then shoot a puppy and call me a slur

I know it's just a figure of speech, but damn. I laughed out aloud in public just reading this.

antonvs 1 day ago||
Kristi Noem would fit the bill, except for the bit about lowering memory prices.
bob1029 1 day ago||||
If we take some time to understand how HBM memory is manufactured (with particular focus on yield risk for final packaging steps), we will hopefully learn that the current capacity crisis is not bullshit.

I guarantee Micron & friends are not intentionally orchestrating their business such that they would suffer a massively reduced chance of yielding on a per-die basis. Unless someone is actually buying HBM devices, they are not going to be making them. These are not a commodity that can be speculatively manufactured in any economically rational way.

javier2 21 hours ago||||
Yeah, I was about to say, the memory industry has been found guilty of price fixing multiple times.
ggeorgovassilis 23 hours ago||||
> Seriously, if a single politician stepped forward and said "i'll bring down ram prices"

Or abolished VAT (the meaning of VAT is that you pay a "rent" for all the infrastructure used to produce the thing) and import taxes (protect your market) on stuff we don't produce in our markets anyway.

fhn 1 day ago||||
How many people would you allow them to kill?
xyzsparetimexyz 1 day ago||||
Neither political party cares at all about memory pieces get real lol
Aerroon 1 day ago|||
I don't really understand why. Memory is a critical component of every computational device.
Exoristos 1 day ago||
That's tautological, but I think you would need to explain how it extends their and their backers' influence to get party attention.
kristopolous 1 day ago|||
wait until holiday shopping...it affects the price of almost everything with a battery or power cord.
xyzsparetimexyz 1 day ago|||
What do you think they'll do? Neither repubs nor dems will touch ai companies in a meaningful way. Anything China does wrt memory fabs week be more significant
kristopolous 1 day ago||
I don't have faith in the political parties. Everything is insane. You look at platter recently? It's up 3x in 12 months, not just ssd or nvme, but straight up traditional platter.

https://web.archive.org/web/20250612003557/https://diskprice...

https://diskprices.com/

After 70 years of decreasing computer prices all of a sudden it's gone 3x, 5x, 10x up in 1 year, we are in total clown world and saying "dur AI" is lazy and doesn't map to reality.

It's Argentina style inflation - as if Honda said "we're only making $500,000 luxury cars now. Everything under $50k we've stopped." and then those cars shoot up to $125k.

It's destroys the market, destroys the consumer, destroys the company, dismantles everything, and they do it for the short term payday.

gchamonlive 1 day ago|||
[flagged]
Analemma_ 1 day ago|||
I don't think RAM vendors have formed a cartel and I think this is knee-jerk anger without any thought. RAM is a commodity product with massive upfront capex costs, and those always have boom-and-bust cycles. At various points in the 2010s and 2020s RAM vendors were getting eaten alive by a supply glut, this would not have happened if they were a cartel.

Is it really so hard to believe that RAM prices are up because demand is simply exceeding supply, especially in a market where additional supply takes years and billions of dollars to come online? There's no need to posit cartel behavior and a fair amount of evidence that there is none.

boustrophedon 1 day ago|||
The RAM vendors have formed cartels previously and been convicted, so although demand is exceeding supply it is not that crazy to at least consider.
kristopolous 1 day ago|||
The AI boom started in 2022. Prices rose THREE years later after 2025 sanction and tariff style legislation to protect the market during a price hike.

I got a 4090 in 2023 for 1600, a 5090 in 2025 for 2000 with 256 DDR5 for about $1,000 ... and then, after some protectionist legislation passed, these prices quickly shot to the moon.

Connect the dots.

Analemma_ 1 day ago||
Man I think you're just spewing word salad and a lot of what you've written is either wrong or not even wrong. The AI hype really got started in 2022, but hype on social media doesn't mean anything for RAM prices, only real buildouts do that. They rose pretty steadily until OpenAI revealed their shenanigans re: locking up a ton of supply from two different vendors with secret contracts, and that's when the takeoff really started. This is definitely scummy behavior from OpenAI (big surprise), and I actually think they arguably should see an antitrust investigation for that (not that that will ever happen), but OpenAI is a buyer; that's not the same thing as the vendors forming a cartel.

You can't say "connect the dots" at the end of a raving, mostly-incorrect post and act like you've made an ironclad argument.

kristopolous 1 day ago|||
There's nothing raving about it. I was trying to communicate that I've got all the hardware I want. I'm still pissed these companies are ripping people off.

It was supposed to be over by now and then they said 2027, then it's 2028, and now I hear "oh it's going to continue to rise the rest of the decade".

I'm likely going to be flying into Shenzhen to put my next computer together. The one I put together in 2025 would have cost me about $25,000 right now. I paid under $5,000.

I can round trip to China for $750. So once they ramp up production that's the strategy.

Other countries already do this. Apple and Google aren't in every country and those people buy new electronics when they travel.

The USA is soon to be on that list

You aren't engaging in good faith and there's no reason to continue with you.

hn_acc1 1 day ago|||
I mean, just because OpenAI started it doesn't mean the vendors didn't form a cartel afterwards (or conspire together) to ensure maximum profits in a "crazy high demand" situation..
petu 1 day ago|||
There's no BF16, original full quality weights are quantized already and 510GB.

Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.

Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).

wren6991 1 day ago|||
Are you counting the n-gram/PLE as part of the model weights there? They can go in host memory. Would be good to show your working. Also the released weights are pre-quantised and presumably QATed, so your "Full Precision" and INT8 are simply not a version of the model that actually exists.

Edit: I went and checked for you. The LM backbone is 307.2 GB (286.1 GiB), straight from DeepSeek's upload. The n-gram table is 203.1 GB (189.1 GiB), which goes in host RAM. Note the embeddings are higher precision than the expert tensors, so it's a larger fraction of the bytes than it is of the parameters.

So,

> Call me crazy but:

You're crazy. :-)

mrinterweb 1 day ago|||
Projects like DwarfStar https://github.com/antirez/ds4 really lower the hardware bar a lot so Deepseek 4.1 flash and other mixture of expert models can run on consumer hardware. There are also other inference providers who make their money serving openweight models. Services like OpenRouter make it all too easy to utilize these models. Access to these models isn't hard. The hardware moat is becoming pretty easy to bridge.
contingencies 1 day ago||
More concretely DwarfStar M5 128GB Deepseek 4.1 flash 1K tokens @ 29s, 5K tokens + reasoning @ 147s, 10k token prompt @ 463 tokens/s = 22s. Hardware buy-in USD$7K / AUD$8.5K / EUR€6.8K. At typical workloads, ROI is still poor vs. current-era subsidies, but owning hardware is good for privacy/longevity/connectivity independence. Whether you actually consider Apple hardware 'owned' is a valid and thought provoking question.
onlyrealcuzzo 1 day ago||
Still gonna take 2-3 years to get DeepSeek V4.1 Flash quality at decent speeds on reasonably priced hardware.

Hardware update cycles are 2-3 years even on the high end, so it's still a ways away before "good enough" and "local" belong in the same sentence for the average person.

And by then, DeepSeek V6 Flash will be too cheap to meter, 5x faster, and 10x better, so... You'd still need to go out of your way.

Most people are spending most of their time on their phones anyway. ..

bitexploder 1 day ago|||
Flash Next is a basically there. It really depends on what you are doing. This model is great. People forget that they felt Opus 4.6 was a great model and now you have it at home.
epolanski 13 hours ago||
DS 4.1 flash is much much better than Opus 4.6, even quantized.
bitexploder 8 hours ago||
I do use DS 4.1 flash a good bit. It's only real fault is that it is a very eager model and I have to system prompt it down to a more reasonable level and it will still take initiative, a little more than I would prefer.
contingencies 1 day ago|||
In the words of a Scottish comedian, "The average person is Chinese." https://youtu.be/LsEhNMy8svo?t=98
lisplist 1 day ago|||
DeepSeek V4.1 Flash is mixed MXFP4/MXFP8 so all but the INT4 calculation is wrong here and that's still wrong because you can run it on < 400GB VRAM. The n-gram table is MXFP8, but can be offloaded to RAM or disk without too much of a performance hit. Really, you could probably cram it onto < 300GB VRAM if you're willing to apply a small quant to certain parts of the model considering how little VRAM is dedicated to kv cache.

I know my comment is a little nit picky because it's still pretty expensive to run, but it's not quite as bad as this comment makes it out to be. Really, if you're VRAM constrained, take a look at GLM 5.3 Flash or Qwen 3.8 Flash Next before you worry about this model as all three models perform pretty similarly.

ct520 1 day ago|||
"1070s or 1070 TIs because GPUs have been severely overpriced for too long" ... ."

1070ti launch MSRP was $450 ish. 5070 could be had in the last year for 5xx-6xx range easily.

All things considered - (inflation being about 30%~ (guess)) between these two timelines. You are looking at 300% performance difference at a cost dollar for dollar that is cheaper then when they purchased their cards.

Might be a bit of a stretch blaming it on "severely overpriced for too long..."

mrheosuper 1 day ago||
The 5070 honestly feel more like xx60ti class instead of xx70 class
ericd 1 day ago|||
In nvfp4, it's about 300 gigs once you offload n-grams, 491 without offloading, you can run it pretty well on 4x DGX Sparks, which last I checked was about $20k. So, it's definitely runnable.

Or you can just use any of the neoclouds' shared hosting. The thing for them to be freaked out is that these models are getting good enough very quickly, and all the shared hosting providers can run them for a tiny fraction of what the frontier model companies charge.

girvo 1 day ago|||
Not quite: not all of this needs to be in VRAM

It has a set of n-gram tables which you can stream from system RAM or even NVMe

That said it’s still quite big! I can’t fit it on my DGX Spark, though I believe you can if you have two?

jonsoft 1 day ago|||
It needs 3-4 Sparks to run well (at an acceptable quantization and sufficient KV cache):

https://github.com/christopherowen/spark-ds41f

girvo 1 day ago||
Ah that’s a shame. GLM 5.3 Flash is honestly as good IMO and can run on two pretty successfully from what I understand.

I’m quite spoiled with how good Qwen 3.8 Flash Next is on a single spark though: shocking how good local models are getting on attainable-ish hardware

jonsoft 1 day ago|||
DeepSeek V4 Flash runs well on two Sparks, I documented that here:

https://blog.jonathanpage.com/

GLM 5.3 Flash runs fine on two Sparks and Qwen 3.8 Flash Next on one is indeed incredible! I made this 3D game with it in two days using Qwen Code as agent:

https://games.jonathanpage.com/

bitexploder 22 hours ago||||
Having Flash Next local at 150 t/s with 250k context is a joy. It’s as good as Sonnet 5. It will spaz out but it was less eager compared to DS Flash 4.1. Both are good but I find I prefer Flash Next. This and Qwen 27B are the models people should be freaking out about.
jonsoft 1 day ago|||
I forgot to say, DeepSeek V4.1 Flash runs great on two DGX Station GB300s!

https://www.storagereview.com/review/dgx-station-gb300-clust...

rsolva 1 day ago|||
I have access to two and will explore this the coming weeks.
girvo 1 day ago||
Also give GLM 5.3 Flash a try: it’s shockingly good too in my testing, and I believe eugr has a TP=2 recipe to use for sparkrun
keammo1 1 day ago|||
The article isn't just about running locally though. The author is saying it's super cheap to run the model through Opencode Go (and presumably OpenRouter etc.) Personally I'm always most excited by models I can actually run locally, but even these huge open source models open up the competitive landscape for companies to let you call models via an API or just lease compute. And they don't have to charge you to offset research, training, huge staffs of the best minds in the world, crazy PR etc. I think that's a big win for customers and buts competitive pressure on the frontier labs as well.
ManuelKiessling 1 day ago|||
Thanks for the data!

Allow a question from someone who’s only got a very vague idea of how this kind of stuff works behind the scenes: say I rent usage of this model through one of the many LLM hosting providers out there, and let‘s assume I use it extensively through something like Pi or OpenCode and vibe code away all the time, keeping the hosted model occupied as much as I can, happily burning my credits.

Does that mean that there is a hardware cluster as described by you above that is crunching away just for me?

So at FP16, I alone keep a 1,664 GiB system occupied all the time?

DrammBA 1 day ago|||
No, a cluster can server multiple users at the same time, providers cap the tok/s so that one cluster can run inference on multiple inputs at the same time. OpenAI with their new ultrafast mode is probably reserving the whole cluster or prioritizing requests of ultrafast users above others with a higher tok/s hence the high price and high speed. There's many other knobs providers tweak that they don't show the users, for example I doubt many providers are hosting the full FP16 version.
jiggawatts 1 day ago||
It's not based on rate limiting at all.

The "expensive part" of generating the next token is streaming in the model weights from memory. The computations are relatively simple, which is called a "low arithmetic intensity" in industry jargon.

So what they do is batch multiple chats together and compute the neuron activations for all of them together.

This is vaguely similar to how some database engines work, where if multiple users need to run a "whole table scan" query, the additional users "join" the streaming workload of the first query mid-way, then loop back around to complete the first part that they missed. The AI accelerators don't do this looping, but the concept is the same: amortize the expensive I/O over multiple computations running in parallel.

The "turbo mode" token rate thing is almost certainly your query getting sent to slower or faster hardware, like B200 vs newer B300 kit.

rnxrx 1 day ago|||
It depends hugely on what "rent usage of this model through one of the many LLM hosting providers" means. If you're asking them to host the model privately then yes, all of that 1.6T of RAM is likely in use holding weights, activations and KV cache by an inference engine that's only getting/answering requests from you alone. When you aren't actively using the model the hosting process is still active and waiting with all of that memory still wired to it.

As background: For the most part VRAM oversubscription/paging/swapping isn't a thing in the same way that RAM for a VM often is. There are some approaches to it, but (to my knowledge) not at that sort of scale.

There are some systemic reasons for this, but very broadly speaking the GPU vendors are building toward the highest bandwidth and lowest latency possible, and the overhead/complexity of something like protected memory modes serves neither of those priorities.

apitman 1 day ago|||
> Even so why would anyone not sleep on a model they cannot run?

Because it's an open model so providers compete on price.

ByteAtATime 1 day ago|||
I think, considering the size of this model, it's closer to a Pro than a Flash on everything other than speed
crossroadsguy 1 day ago|||
I did somet math and completely gave up on the idea of trying any worthwhile local model and figured I'd rather pay the 15-30 USD per month via subscription and/or API key combos for years than buying a local setup which might go out of date very fast, if it doesn't goes kaput just out of warranty. I won't be surprised if RAM scarcity is an concerted effort to herd people towards the remote models :)
cookiengineer 1 day ago|||
It's dangerous to go alone. Take this: [1]

I reimplemented most of the features of the Deepseek v4.1 flash paper (apart from quantization aware training which doesn't make sense because my implementation uses float32 precision anyways)

I'm currently learning how to distill reasoning traces (check my other github repositories) but I think that a locally selfhostable deepseek is possible with my mixture of experts sharding mechanism. I decided to optimize everything for CPU parallelization, with the idea that the KV cache and meta model have to run from CPU RAM anyways, so the experts can also be loaded/unloaded at runtime if needbe, to save more RAM.

My assumption is that the KV cache optimizations in combination with the CED and compressed attention features are the reason why v4.1 flash has so few hallucination problems and such a strong self-lookup/thinking behavior. But that's more a gut feeling, need to evaluate and test this more thoroughly.

Anyways, would love to see someone train this on their own datasets. Currently my pipeline is kinda optimized for parquet and zim files.

[1] https://github.com/cookiengineer/gonano

ls_stats 11 hours ago|||
It's funny because VRAM IS cheap to produce.
nicman23 22 hours ago|||
what they should freak out about is qwen 3.8 github.com/Niko1221/Strata

it rips with just 64 ram and a 9070xt

sgt 19 hours ago||
What would be the ideal qwen 3.8 variation for a 5090 with 32GB?
nicman23 18 hours ago||
q8 probably with the link above
anvuong 1 day ago|||
I just un-retire my pair of 1080Ti for some small models development because the current GPU prices literally make me sad.
CamperBob2 1 day ago|||
You can run it locally for the price of a decent car, or run it (hopefully) privately on somebody else's hardware at vast.ai or a similar provider for much less. What's not to like?

No, you won't get frontier-level intelligence on a 1070Ti. Yes, it should be illegal to do what Altman did. Since we clearly don't live in the best of all possible worlds, we need to settle, and DS4.1 Flash is a good place to do that.

For tasks that don't require vision I personally like the NVFP4 quant of GLM 5.3 from Local Inference Lab better than DS4.1F, but they are both well beyond awesome.

nullc 1 day ago|||
The bulk of its weights are natively MXFP4. And engram values don't need to be in vram.
functionmouse 1 day ago|||
one can make a fine gaming pc for ~$350

1660 ti, 4790k, 16gb ddr3

dgellow 17 hours ago|||
Look at https://commandcode.ai/pricing, you can have very decent volume of requests with deepseek flash v4.1 for literally $1/month
holoduke 1 day ago|||
He doesn't ruin the cost of memory. Advances in memory size and speed are now in full speed mode. Expect drastic increase in the upcoming years. Big factories are in the making and planned. Gigalab in the US and many others in the east. Since 2010 we have computers with 16gb as being normal. Finally we are moving into a new era where the standard will be 64gb next year and 128 in 2028. Hopefully we reach 1tb in 2030.
CorrectHorseBat 1 day ago||
I've read the exact opposite, vendors are reducing the standard from 16GB back to 8GB
holoduke 1 day ago||
That's only temporary till production meets demand again.
CorrectHorseBat 1 day ago|||
Which is not going to happen in the upcoming years
Tepix 18 hours ago|||
Remind me in 4 years then? Or 6? or 10?
liuliu 1 day ago|||
What are you talking about? The model is native NVFP4, why you run it at any precision higher than that?
jauer 1 day ago||
This “blame sama for memory prices” meme is so tired.

He gave demand signal so many times years ago and was mocked for it and now we have the consequences of industry not taking him seriously.

More comments...