Top
Best
New

Posted by dnhkng 17 hours ago

DeepSeek-V4-Flash Update(api-docs.deepseek.com)
660 points | 315 commentspage 4
vladukha 15 hours ago|
Where do you guys get deepseek? I'm hearing a lot of good reviews and want to try it with my pi config. from the deeepseek themselves, openrouter, or anywhere else? does it make a difference? [edit]: whoa it is really fast. will take some time to evaluate quality thou
u8080 13 hours ago||
Directly here: https://platform.deepseek.com/ Easy top-up and pay as you go. Availability and speed are very good and it is the cheaper option.
psibi 14 hours ago|||
I've been using DeepSeek directly. I've heard from colleagues that using it via OpenRouter is slower, but I'm not so sure about that.
ticoombs 14 hours ago|||
> does it make a difference

Probably not.

But Opencode-Go is a great solution for those who don't want to pay DeepSeek directly (or can't due to reasons)

Selfish referral code: https://opencode.ai/go?ref=R1AJZT4VBX

embedding-shape 14 hours ago|||
For hosted APIs, it's a lot cheaper to use their own infra, caching seems a hell of a lot better there compared to OpenRouter, and indeed the tok/s seems higher. They also have peak/off-peak pricing, so if you can hold off with your request, you get a pretty big discount.

Otherwise, if you're trying to run it locally, even really low quantizations like DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix seem to actually not be so dumb compared to smaller models with same quantization, might be worth a try if you're sitting on a lot of RAM/VRAM yet not industry-scale amount :)

chronogram 12 hours ago|||
I use it directly: https://platform.deepseek.com/usage

3rd party providers on OpenRouter can be cheaper but it's already so cheap.

kzrdude 13 hours ago|||
Opencode-go gives you $60 worth of DS V4 api usage for $10 per month. Right now I think it's hard to exhaust that when using flash exclusively, and plain API use might even be cheaper! Anyway, for DS usage it's a good deal.
gpugreg 12 hours ago|||
To add to this, the $60 only applies to DeepSeek-V4-Flash and a few other models. For DeepSeek-V4-Pro, the amount is $15.

https://opencode.ai/docs/go/#usage-limits

Previously, OpenCode Go had higher API prices for some models, but now they lowered the API price and simultaneously reduced the allowance.

kzrdude 12 hours ago||
Thanks. These things change day by day I guess, AI is just moving fast (and I'm on vacation).

GPT 5.6 Luna is a new model in Go since I last checked, for example.

Lalabadie 12 hours ago|||
Opencode also have a ZDR (zero data retention) deal with them – if I recall correctly, that's not something you can enable as an individual DeepSeek subscriber.
anon373839 12 hours ago|||
I really like OpenCode - BUT: there is a loophole the size of Portugal in that ZDR language. All they say is that their providers follow a ZDR policy. I haven’t found anything promising that OpenCode themselves don’t retain Go usage data. Something to be mindful of.
gpugreg 12 hours ago|||
Unfortunately, all mentions of ZDR have silently been removed from the OpenCode Go page today.
kzrdude 12 hours ago||
Thanks, any update here is important. I use them because of good data policies..

I still find this today:

> The plan is designed primarily for international users and provides stable global access. Your data will not be used for model training.

gpugreg 9 hours ago|||
dax (coauthor) recently tweeted https://xcancel.com/thdxr/status/2083178051052155182

> because we added the new deepseek which we do not yet have a ZDR with we cannot blanket say we offer ZDR

I wonder how the website can make the statement that data will not be used for training.

Tepix 10 hours ago|||
Do they do other things with the data, like selling it?
lucianmarin 11 hours ago|||
OpenCode harness gets the most out of DS V4 Flash model. You can implement any coding task, fast and cheap.
Gigachad 14 hours ago|||
I used it through openrouter. Plugged in to the vs code copilot bring your own key thing.

Played around for a few hours and used up 80 cents of tokens.

Lalabadie 12 hours ago||
Anecdotal data from my own tests: allow only one provider if you want good cache usage on OR. 80 cents is probably 4x the price you should have paid.

Providers' cache hit stats are available to consult, and only 1-2 of them behave properly if I remember correctly, zero if you request providers that don't store and train on sessions.

darkest_ruby 13 hours ago||
Openrouter
egeozcan 15 hours ago||
Every time I want to have fun coding something with natural language processing, I use deepseek flash. It's just incredible for the price. I have a fairly popular app with 400 users that uses DeepSeek in the background and it still didn't hit even 50 bucks of usage in a month.
indigodaddy 11 hours ago|
Mind sharing the name? Sounds interesting.
HyperL0gi 11 hours ago||
Is anyone using DSv4 for their agents that is not related to writing code? Curious about use cases specially for someone using gpt-5.4 mini for classification, categorization, etc
freakynit 10 hours ago||
I use it to conduct thorough online researches. Plug-in some online search MCP (like the one I use: https://jerrysniffs.online ), and the flash models dig through the internet for dirt cheap..

Most of the times, the total cost, including search API's, is less than $0.05 for full deeply researched output, and the research is actually good.

Lalabadie 9 hours ago|||
It's been good for one-off cases in my limited experience. I would describe its behaviour as Sonnet-shaped, if that makes sense to you. Good answers but it often decides to reason a lot about simple things before getting to an output.

At the speed Flash has on most providers, it doesn't really turn into a latency concern.

markab21 9 hours ago||
We use it at a moderate scale, self-hosted on B300 hardware. It's great :D

QA analysis of voice transcriptions. Napkin math: we operate at 2-5% of the cost of running on Equiv Frontier, though this changes near-weekly because pricing is so volatile.

It took us about a month to get the inference configured to achieve these numbers. But if you can get your hands on a pair of B300 GPUs and the context works, it's untouchable for price/performance.

(B200 would work, but you don't have the B300's memory, which lets you run it on 2xGPU instead of 4xGPU... with Dspark, it's like magic)

On a side note, for tasks that don't require the intelligence of DS v4 flash, we're using Nemotron-3-super with incredible success. I'm shocked we're not seeing more adoption of this model, given how easy it is to fine-tune and how blisteringly fast the nvfp4 version is. (A single B200 GPU can produce an insane amount of throughput with Nemotron 3 Super.)

PhilippGille 15 hours ago||
The previous V4 version wasn't called “Preview” by most inference providers. For example, the OpenRouter model slug was `deepseek/deepseek-v4-flash`. So now there will be confusion when someone talks about V4 Flash or when someone offers V4 Flash inference.

Why not call it V4.1?

petu 15 hours ago||
It probably would be called 'deepseek-v4-flash-0731' in API

edit: nope, at least deepseek kept "deepseek-v4-flash" and just updated model underneath. I guess preview is no longer worth serving with that release and you'd have to look through inference provider docs to see if they've updated, yeah..

PhilippGille 13 hours ago|||
That's what I mean. On DeepSeek it's now just `deepseek-v4-flash`, while OpenRouter calls it `deepseek/deepseek-v4-flash-0731`, so now when someone talks about DeepSeek V4 Flash, like in benchmarks, or other inference providers, which version do they actually mean?

The `-0731` style suffix is worse compared to a proper version bump like V4.1.

Macuyiko 12 hours ago||
What I and my team have been doing in papers is just to refer to the OpenRouter slugs. It upsets reviewers because they will complain it "is not sufficiently clear to a wider audience" but I do agree it's the cleanest approach. Also goes to prove how much power OpenRouter has actually...
kzrdude 13 hours ago|||
Does it still say that it's an anthropic model, when asked? I would guess new post-training has fixed that.
petu 13 hours ago||
Does Claude still say it's Deepseek, when asked?

https://news.ycombinator.com/item?id=49082022#49087112

How is that important? Maybe it does, so what?

kzrdude 12 hours ago||
Not important, just a curiosity. I would expect that kind of knowledge would be pretty well burned into the weights, but what do I know about LLMs. Gemma 4 can answer this question well.
bermudi 13 hours ago|||
DeepSeek being DeepSeek. v3 and R1 went over the same and had multiple versions
try-working 15 hours ago||
DeepSeek themselves called it `deepseek/deepseek-v4-flash`. Pro is still like that.
PhilippGille 13 hours ago||
Yes that's my point. The old and the new version are different in capabilities, but now when someone talks about DeepSeek V4 Flash (in benchmarks, on inference providers), you don't know which exact version it's about.

Some providers like OpenRouter now call it `deepseek-v4-flash-0731`, but even in places like here on HackerNews people say things like "Sonnet is better than DeepSeek" without specifying a version or a reasoning effort, certainly no one will mention that `-0731` suffix when talking about DeepSeek V4 Flash.

benjiro29 12 hours ago||
I think it does not matter. Most people who use these types of models are into the IT world, and will know/be informed very fast that there is a difference. People will likely also use the 0731 behind it.

The only issue i see, is 3th party providers that have not yet updated. But that is going to be a short time periode. There is no reason to not update.

flysoft 15 hours ago||
Finally have a model with usable intelligence, at a reasonable price. Can't imagine what Pro GA would look like, considering pro preview has only 1.6t parameters.
throwaw12 14 hours ago||
how different is their harness from Pi coding agent harness, is it possible to make an extension for Pi which can implement deepseek harness?
storywatch 15 hours ago||
How's their performance in English prose? We are currently searching for cost effective ways to keep story wikis up to date.
amunozo 15 hours ago|
I am interested in this too, as I think it could be these models are overly optimized for coding. Let me know if you figure it out!
w2seraph 4 hours ago||
This made my day !
nathaah3 14 hours ago||
DS v4 flash has been my goto model for tasks in work. its been unsurprisingly fast and cheap.
k__ 13 hours ago|
I'd take more throughput while everything else stays the same.
More comments...