Top
Best
New

Posted by theanonymousone 15 hours ago

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis(artificialanalysis.ai)
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
503 points | 277 commentspage 4
nilsbunger 8 hours ago|
Is artificial analysis using the 80% reduction in Luna pricing that was announced yesterday in these charts?
hxii 10 hours ago||
I’m wondering if they did anything to address the DSML tool calls leaking. Has been an issue with both Flash and Pro so far.
denismi 9 hours ago|
On OpenRouter it seemed to only affect a couple of specific providers. Ignoring them has made it s non issue for me.
smrtinsert 7 hours ago||
I have been loving deepseek flash. I can code in my hitl preference doing complex but moderate amount of coding and not even hit 1 dollar in a session. It gets harder to justify paying 20/mo to Anthropic for unused compute when I have it on demand and at a much cheaper rate with deepseek.
paoliniluis 11 hours ago||
Would be awesome to see a new ds4 release. Having so much in something that can be run locally is mind blowing
k1e 11 hours ago||
No speed (tokens/s) benchmarks?
k__ 11 hours ago|
On OpenRouter it's 93 TPS.
prism56 2 hours ago||
Need a few more providers that wont violate ZDR and i'll switch over.
sourcecodeplz 3 hours ago||
from my testing it is at glm-5.2 levels
tmikaeld 9 hours ago||
My problem with DS flash/pro is that they don’t push back on obvious bullshit, both irl and code [0] but it’s a great implementer workhorse if you give it _very_ detailed specs.

[0] https://petergpt.github.io/bullshit-benchmark/viewer/index.v...

net01 10 hours ago||
here it is on openrouter https://openrouter.ai/deepseek/deepseek-v4-flash-0731
k__ 9 hours ago|
Half OT:

Why do the cache hit rates seem to vary so much between harnesses?

I use pi, which is very minimalist, and I get a hit rate of ~99%. Paying like $1 a day for Flash. Yet, the hit rate mentioned on OpenRouter is only ~79%.

drob518 9 hours ago||
Yes hit rate does vary by harness and by how you use the harness. If you use subagents, for instance, they will start with a whole new context created by the main agent, and this will not be cached. If you mostly use the main agent with Pi, you’ll have high hit rates and low costs. Sometimes agents do “cache busting” things where they’ll move around some of the text in the context to try to keep old instructions from being forgotten, thus keeping the agent on task, and this will bust the cache. I’ve heard, but not validated myself, that Open Code has some issues with this.

BTW, this is one of the things that I really like about Pi. It’s very simple and thus very predictable.

hnsmomdpvp 7 hours ago||
The details are where it gets interesting
Aeroi 6 hours ago|
so are we all just coding for free now?

create a plan with SOTA, execute with this.

More comments...