Top
Best
New

Posted by dnhkng 18 hours ago

DeepSeek-V4-Flash Update(api-docs.deepseek.com)
665 points | 319 commentspage 5
kamikazechaser 16 hours ago|
The flash variant is on par with Sonnet 5 on DeepSWE (54%). Big, if true.
spwa4 18 hours ago||
In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.
wolttam 18 hours ago||
It runs really well on 2 DGX Sparks - 60t/s
Tepix 11 hours ago||
Yes, the Dual DGX Spark looks like the sweet spot for this model for now. Good preprocessing speed. Lots of context. Fast enough for 1-5 devs perhaps. Around 8200€ as of today (used to be 6000€).

Dual Strix Halo is much slower and current Macs with 256GB are both slower and more expensive (Mac Studio M3 Ultra 256GB around 12000€).

To get something faster than the two Sparks you'd need to spend more than $22000 for a server with 2x RTX Pro 6000 at $10000 each.

Beyond that you could get 2x AMD MI350P.

arjie 16 hours ago|||
Not yet, right? That's the old DS V4 preview release. We're still waiting for the weights to come out.
benjiro29 13 hours ago||
Probably the same. When the same base model is trained, the weight do not tend to change a lot. GLM 5.0 > 5.1 > 5.2 are the same base model, that just kept being trained. Weights hardly change as a result. Think in the like few percentage points size difference.
lukan 16 hours ago||
"it'll barely run on an M5 Max "

The max version I could order now with 128 GB?

If so, the price for local inference would be 12 000 € vs 500 000 € for a B300.

NitpickLawyer 16 hours ago|||
There's also the 2x spark way, which should be ~8k eur? Someone down the thread reported ~60tps for 2x sparks. That's totally usable for local inference.

You can also do 2x 6kPRO in a workstation, for ~20k.

matrik 16 hours ago|||
For the same performance, one could even go about 50% cheaper with 16 channel ddr4 + a rtx3090 for prompt processing.

But still, even for mid level projects API is orders of magnitude cheaper, since you don't need to set it up and maintain it.

gpugreg 13 hours ago||
The memory bandwidth of the 2x RTX Pro 6000 Blackwell setup will be 10x higher, which should have an equivalent effect on the generated tokens per second.
spwa4 16 hours ago|||
Currently the 3bit (and 2 bit) quant on DGX spark (on one of them) and the M5 Max should just start. Right now. (I'm hoping to get an M5 Max delivered on monday, let's see if it happens this time. It's 2+ months since I ordered now)

The 4 bit quant technically fits (there's a 127 GB version) but ... obviously that's not going to work. It is so close though, surely someone will a way to do it.

reverius42 16 hours ago||||
I'm running a useful quantization of the previous version of Deepseek-V4-Flash -- quite well but with so much fan noise -- on a MacBook Pro M5 Max with 128 GB.
spwa4 16 hours ago|||
500k is for the 8x B300 version. Which is the only one you can buy atm. But technically a B300 card is more like 60k, just impossible to get.
znnajdla 14 hours ago||
The conspiracy theorist in me wants to think that the 80% drop in GPT 5.6 Luna prices today is correlated with this update from DeepSeek. Perhaps OpenAI has already hacked its competitors with it's Mythos-like models and is aware of what competitors are doing and is able to react in advance.
nchmy 12 hours ago|
Seems to me that it's the reverse. Open ai dropped prices to compete with deepseek etc and then deepseek released this update to negate Luna.
ra 16 hours ago||
What's the best way to run this on a 64GB M2 Pro?
petu 15 hours ago||
Weights are yet to be released (maybe in 24H, Deepseek has track record of releasing same day).

https://github.com/antirez/ds4 is often mentioned for DS4F on Mac, but 64GB is likely not enough to achieve reasonable speeds (official weights should be ~160GB).

Tepix 11 hours ago||
DS4Flash has 284B weights. 64GB? No go.
zozbot234 10 hours ago||
The native weights are 4-bit for the sparse experts, and they quantize to ~80GB with limited degradation in real-world performance. That's a viable target for 64GB with SSD streaming, though it will be slower than keeping the whole thing in RAM (especially on a M2-class machine with its slower storage).
XCSme 12 hours ago||
Can't really use it now, without giving away your data:

> Trains: this provider may use prompts for training and may retain prompt data.

Philpax 11 hours ago|
That will cease to be a problem in the next 24 hours, now that the weights are out: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
sparse-Matrix 14 hours ago||
This may come as a surprise to a lot of AI concerns, but I have -zero- interest in paying for a model.
kzrdude 13 hours ago||
Then DS V4 flash is pretty relevant, because it's pushing prices down.. as well as being self hostable with a large enough rig.
Tepix 13 hours ago||
Why mention it? Just download the weights when they become available.
sreekanth850 11 hours ago||
how this compare to luna high with reduced pricing.
mekky16 14 hours ago||
if they were anthropic they wouldve just released it as a new model
truth_seeker 14 hours ago|
The magic of post training with valuable dataset
More comments...