Top
Best
New

Posted by theanonymousone 16 hours ago

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis(artificialanalysis.ai)
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
510 points | 281 commentspage 5
Aeroi 7 hours ago|
so are we all just coding for free now?

create a plan with SOTA, execute with this.

lostmsu 15 hours ago||
https://artificialanalysis.ai/models/deepseek-v4-flash
buildinext 9 hours ago||
are folks checking latency for applied voice for newer models asa standard?
epolanski 12 hours ago||
I was writing a benchmark for my own harness, and DS4 flash answers as well as Fable 5 on any query.

The specific agent is focused on getting precise and on point answers about a codebase.

The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all.

The benchmark included more than 50 questions or different difficulty.

But when the agent was improved in its prompt and rooting it was impossible to have it perform worse than closed source sota.

Just to say that the quality of the harness is as important as agents intelligence.

quikoa 3 hours ago|
What does your harness do to squeeze that value out of DS4 flash if you don't mind sharing? It'd interesting how it compares to other harnesses (even if it's a qualitative assessment instead of quantitative).
freakynit 12 hours ago||
I will get downvoted, but fck it.

The ban on these open models is coming within weeks, if not days. As usual, the excuse will be "national security".

UltraSane 12 hours ago||
How exactly will they ban them?
freakynit 12 hours ago||
By making companies using them "toxic" to touch.

For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models.

They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for such open models.

qphe95 10 hours ago|||
Can't wait to distill Deepseek v4 flash to America-1
nancyminusone 10 hours ago||||
Some contractors are already barred from using Claude due to DoD designation back in March as a "supply chain risk"
tyfon 12 hours ago||||
So now the US companies will be stuck on expensive models while the rest of the world can do things much more cost effective.

I'm not sure the outcome would be beneficial for the US as a whole here. But perhaps that is not their priority.

freakynit 10 hours ago||
There are many proxy-ways to bypass such restrictions. Of course, the costs will be higher. Providers will spun-up.
UltraSane 10 hours ago|||
Companies could self-host models in secret. It would be hard to stop.
freakynit 10 hours ago|||
Not legally. Liability is a big thing.
dgellow 10 hours ago|||
The thing is, you don’t need to actually block usage to make something illegal. You make it so toxic that company wants to be seen publicly using open models
UltraSane 10 hours ago||
Companies would secretly use self-hosted models internally because it would give them a enormous cost advantage.
dgellow 9 hours ago||
Sure, companies can indeed do illegal things
UltraSane 8 hours ago||
But how would the actual weights it be made illegal without violating the 1st amendment?
hgoel 8 hours ago||
By ignoring the 1st amendment? The US has a tendency of ignoring its constitution whenever it's convenient.
UltraSane 7 hours ago||
Don't be glib, the 1st amendment is taken VERY seriously.
dgellow 5 hours ago||
Im not sure if you’re joking or not.

- the US strictly regulate cryptography https://en.wikipedia.org/wiki/Export_of_cryptography_from_th...

- some prime numbers are considered illegal https://en.wikipedia.org/wiki/Illegal_number#Illegal_primes

I don’t think banning open weight is that far fetch compared to those

Der_Einzige 11 hours ago|||
Hell, the US doesn't even need to act.

I claim the CCP will wise up within 2 years, possibly much much sooner, and ban their own companies from open sourcing to prevent the Americans from acquiring the capabilities.

Despite all the nonsense claims of China distilling US models, the reality is that the Americans absolutely do distill these free Chinese models, and distillation when full logprobs are available (i.e. you have access to the weights of the model) is an order of magnitude better than when you don't.

Yes, Chinese open weight models in the short term harm US closed source model providers bottom line. In the slightly longer term, "showing your hand" and publishing both the architecture innovations and the models weights will be too dangerous for the CCP to allow. This is triply true if they can release a model that beats the Americans on most benchmarks.

I've already warned investors that this is probably the closest open weight models will ever get to closed access.

VulgarExigency 11 hours ago|||
They may reverse course in the future, but the current directive from the CPC is that Chinese AI labs should be open sourcing their models.

https://www.businessinsider.com/xi-jinping-open-source-ai-us...

mcbuilder 10 hours ago||||
I don't think it's necessarily "wiser" to go closed source. All of AI is built on mostly openness, at least on the software side. There are other ways to compete, it's just the model itself will be a commodity.
rapatel0 9 hours ago|||
You’re describing the dump and pump strategy that china has historically used across a number of industries. Stands to reason that this is what is going on.
dgellow 11 hours ago||
why would you get downvoted, that's one of the most obvious next step
Der_Einzige 11 hours ago||
Because reddit unironically has better decorum around usage of their upvote/downvote system than HN does.

People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!

nickthegreek 10 hours ago|||
You define OPs post as "objectively correct information" even though it is an unknown future event for which they provided zero evidence?
freakynit 8 hours ago||
There's nothing wrong about that. Historically I have experienced this behaviour on hackernews multiple times, hence my comment.

Since when have hackernews started to become toxic like stackoverflow used to be?

dgellow 11 hours ago||||
fwiw pg said early on that downvoting for disagreement is perfectly fine: https://news.ycombinator.com/item?id=117171

commenting about voting is also something the HN guidelines warns against:

> Please don't comment about the voting on comments. It never does any good, and it makes boring reading.

https://news.ycombinator.com/newsguidelines.html

bellowsgulch 8 hours ago||
Reminds me of telling employees to not discuss their wages. If voting here is such a toxic experience, maybe HN needs to wake up and smell the garbage.
NooneAtAll3 12 hours ago||
what a horribly heavy and resource-consuming website...
bigmadshoe 10 hours ago||
It looks awful on mobile too. Barely readable in many parts.
WithinReason 12 hours ago||
we need a benchmark website benchmark
theanonymousone 11 hours ago||
A benchmark website to benchmark benchmark websites?

Or a benchmark to benchmark benchmarks?

NooneAtAll3 10 hours ago||
website that benchmarks benchmarks is benchmark benchmark website

benchmark website benchmark is indeed a benchmark that benchmarks websites with benchmarks (but it can be shown outside websites as well, it's not picky)

Computer0 8 hours ago||
Have not had great ds performance in agentic harnesses in the past compared to glm or k3
0xchamin 10 hours ago||
page not available for me.
try-working 12 hours ago|
Now let's see Dario's price cut.
_ache_ 8 hours ago|
I think, it's the end of Anthropic.

They can't compete, they have bills. By the end of the year, if they can't react, it's game over.

Maybe US clients could be a little patriotic here, but money is money. They won't give them free money forever.

try-working 3 hours ago||
yep. I've been saying for month that exactly this will happen with DeepSeek and Kimi K3, and margins will erode, and Anthropic will fail to IPO.
More comments...