Top
Best
New

Posted by tosh 11 hours ago

DeepSeek V4 Flash 0731(arcprize.org)
534 points | 316 commentspage 3
surprisetalk 11 hours ago|
This reminds me of those pareto-style speedrun record charts when a new glitch is discovered.

[0] https://taylor.town/silver-landmines

When I see dramatic leaps like this, it tells me that the important hacks haven't yet been discovered.

zmmmmm 7 hours ago||
it's great but we need a multi-modal model of this quality and price to truly declare victory.

But it makes me quite curious, how a text-only model can do so well on ARC-AGI-2 being a set of visual puzzles? It would have to solve it entirely using text-only spatial reasoning about the grid (or maybe writing code?). I am curious if this is normal or do other models use their vision capabilities to solve the puzzles?

ghosty141 6 hours ago|
xiaomi mimo is very cheap and not bad.
xyzsparetimexyz 10 hours ago||
That page needs a Pareto frontier display. But wow, it absolutely demolishes.
Almondsetat 6 hours ago||
It's crazy to think V4 Pro still hasn't finished the post processing.
minimaxir 11 hours ago||
It's always fun when Max reasoning is cheaper than High reasoning.
Terretta 10 hours ago|
Rework is expensive.

Tell your PjM who should tell your PgM who should tell your PdM, all the PMs...

Maybe if "the business" sees it is true of LLMs, they might believe it's true of giving better context to engineers up front then giving them time to think and prototype (thinking tokens are an answer prototype).

sourcecodeplz 10 hours ago||
wow. i remember when GPT-5.2 (medium) was everyone's favorite.

ARC-AGI II:

- GPT-5.2 (medium) %26.7 ($0.759)

- DSV4-Flash (max) %61.4 ($0.04)

nmitchko 8 hours ago||
Perhaps it might be interesting: a latent thinking version is here https://huggingface.co/nmitchko/DeepSeek-V4-Flash-0731-Laten...

Does no thinking emissions for context saving.

kamranjon 8 hours ago|
This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?
gentlewater 11 hours ago||
I’ve been refreshing hacker news constantly for a week now waiting for v4 pro, after they stated it would follow «soon». I have learnt «soon» is a matter of definition.
indigodaddy 10 hours ago||
I guess you mean a "new" v4 pro?
ignoramous 10 hours ago||
> been refreshing hacker news constantly for a week now waiting for v4 pro

https://reddit.com/r/DeepSeek is where the fellow F5ers are at.

KolmogorovComp 9 hours ago||
Looking at the caching price of deepseek compared to its competitors, does it have a secret sauce or is it just subsidizing?
ignoramous 8 hours ago|
That's DeepSeek's way of selling "token plans", yes. But without the downsides like daily or weekly limits and guaranteed upfront/fixed spend.
seanmcdirmid 7 hours ago|
Note they double the price if you use during peak time. However, they define peak time with respect to China, not Europe or the USA...so if you are out of Asia, I guess Australians might be impacted, and its still cheap anyways.
More comments...