Top
Best
New

Posted by theanonymousone 14 hours ago

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis(artificialanalysis.ai)
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
493 points | 272 commentspage 3
cmrdporcupine 7 hours ago|
This seems to me like this is probably at least a large part of what OpenAI was up to yesterday with their aggressive price cutting; trying to get out in front of this.

If the full non-flash model follows up with the expected improvements, and at the price point they've been keeping, it puts the frontier labs in a tough position and it feels to me like like OpenAI is reaching deep into their pockets to try to head that off.

TFA link is a 404 though. I'm reading through https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 instead

darknoon 5 hours ago||
Really want to use this, but for web page design you really need vision!
embedding-shape 12 hours ago||
Is the "Output Tokens per Intelligence Index Task" data actually correct or am I reading it wrong? It says there that "Kimi K3 (Max)" would think/reason less than than deepseek-v4-flash, and a whole bunch of other models, like less than hy3 and even gpt-oss-120b, but in my experience, K3 is probably the model that thinks/reasons the longest of all of these.

Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point.

Lwerewolf 6 hours ago|
Just tried the preview on my little test codebase and a "check this out and tell me what you think" prompt used over double the tokens of the previous iteration, but it was a lot more eager as well. Kind of reminds me of the new laguna (s 2.1).
storus 9 hours ago||
I hope they somewhat fixed the hallucination and forgetting plagued V4 previews and that it wasn't just benchmaxxed but the numbers hold in reality. Then it would be my choice for 2x DGX Spark or 2x RTX Pro 6000.
SomeHacker44 10 hours ago||
> DeepSeek V4 Flash 0731 (Reasoning, Max Effort) is amongst the leading models in intelligence and well priced when comparing to other models of similar price.

Similar price? Doesn't make sense. Maybe they meant power, capability or speed?

gorkemyildirim 8 hours ago||
404 ? It's very strange that this can't be fixed.
sim04ful 7 hours ago||
I can't help but feel the timing coinciding with luna's price updates to be somewhat strategic. But without multi-modality it's slightly dead-in-the-water for my usecase: https://design.withfudge.com. I'm currently using Minimax-M3, but Luna ekes out abit futher on the intelligence, so i'll be switching to it very soon.
nilsbunger 7 hours ago||
Is artificial analysis using the 80% reduction in Luna pricing that was announced yesterday in these charts?
mrnobody_ 8 hours ago|
I'm testing it right now and I was genuinely impressed. I've used same tasks I had run the other day against the flash preview.
More comments...