Top
Best
New

Posted by dnhkng 19 hours ago

DeepSeek-V4-Flash Update(api-docs.deepseek.com)
668 points | 320 commentspage 6
sourcecodeplz 16 hours ago|
i've made a comparison between this and GPT Luna (recent %80 price drop)

https://x.com/SourceCodeplz/status/2083099712760987746

i prefer GPT-5.6 Luna honestly

dakolli 15 hours ago||
Chinese labs rushing to release models this week, because it's inevitable that Washington regulates Chinese models in the next 4 weeks. All the US AI leaders have been taking trips to Washington this week, what do you think they're there for..
Tepix 14 hours ago||
[dead]
tosh 18 hours ago||
[dead]
dnhkng 19 hours ago||
DeepSeek V4 Flash (Preview → 2026-07-31)

• Terminal Bench: 56.9 → 82.7 (+25.8)

• Toolathlon: 51.8 → 70.3 (+18.5)

Compared to GPT-5.6 Terra:

• Terminal Bench: Flash 82.7 vs Terra 78.4

• Toolathlon: Flash 70.3 vs Terra 53.1

• DeepSWE: Flash 54.4 vs Terra 69.6

• Agents' Last Exam: Flash 25.2 vs Terra 50.4

Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing scores. Very interesting!

villish 18 hours ago||
> Terminal Bench: Flash 82.7 vs Terra 78.4

Terra 87.4

https://openai.com/index/gpt-5-6/

benjiro29 14 hours ago||
https://www.tbench.ai/leaderboard/terminal-bench/2.1

> 78.4

The real score is always the official benchmark.

We need to see later if DS4 flash 0731 is going to maintain the score but we need to look at the official benchmarks.

Already seen a PR for DeepSWE to update the benchmark with 0731, so we can verify claimed vs official.

villish 10 hours ago||
The comment I replied to was comparing numbers released by each lab, and they made a typo with that specific benchmark.
throwaw12 18 hours ago|||
Open flash model is competing against OpenAI's 'Sonnet' model at the price of GPT 3, I am really excited about this release, hopefully it holds up in real work as well
bayesianbot 17 hours ago||
IIRC GPT 3 was priced at per 1k tokens, had to check, the biggest GPT 3 model from OpenAI was $0.06/1k, so $60 / 1M. gpt-3.5-turbo was the first model after ChatGPT and that was $2 / 1M. And no caching. So not really in the same ballpark
Iolaum 19 hours ago|||
Since they did this with their own harness I m not sure it's apples to apples comparison.
NitpickLawyer 19 hours ago|||
> not sure it's apples to apples comparison.

They're literally comparing the previous version of the same model with the new one. It's based on the same architecture, same pre-trained model, just different post-training. It doesn't get more apples to apples than this.

dnhkng 19 hours ago||
I think the commenter means the Flash vs Terra benchmarks.
NitpickLawyer 17 hours ago||
Ah, my bad. Yeah that makes sense. They do say "The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex.", so at some point someone will make a "same harness" comparison.
dnhkng 19 hours ago|||
It will be fair if they release the harness though. I think now the future will be paired model-harness releases, not just weight dumps.

The performance changes are so big with the right harness that is makes sense to engineer the harness and fine-tune the model to one another from the start.

yms_hi 18 hours ago||
I think it's better than GPT Luna.
try-working 18 hours ago||
Let's see how the market reacts.
Havoc 17 hours ago|
>benchmark results far exceeding V4-Pro-Preview:

Wow that's crazy

Good times for those that don't need strict data protection

blackoil 16 hours ago||
These being open provide much better data protection.
kzrdude 15 hours ago||
Open weights means that there is a free market on providing inference using this model. Some of them will offer good data protection. (Well, hopefully)