Top
Best
New

Posted by dnhkng 14 hours ago

DeepSeek-V4-Flash Update(api-docs.deepseek.com)
638 points | 305 comments
NitpickLawyer 14 hours ago|
This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.

DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.

Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.

dotancohen 3 hours ago||

  > Whatever capabilities they get, can be used "forever" going forward.
"Forever" gets the scare quotes because it is implied only up until the Butlerian Jihad?
sigzero 1 hour ago||
I upvoted you just for the Dune reference.
idiotsecant 1 hour ago||
I feel like once they make a mega successful movie of a book we can stop doing the secret handshake.
trvz 28 minutes ago||
Especially when the movie is over forty years old.
chorizo 5 hours ago|||
Starting to wonder if the free big pickle model on opencode has been DSV4F0731 for the past few months. It’s been incredibly fast and good.
kevincox 4 hours ago||
At least in the past it was GLM-4.6. IDK if it is ever changed.

https://github.com/anomalyco/opencode/issues/4276

chorizo 3 hours ago||
I’ve heard that too, but since it’s a stealth model, I suspect they change it to whatever preview model provider that offers a free endpoint. In the past few months, I’ve gotten 4xx errors identifying the provider as DeepSeek.
dnhkng 14 hours ago|||
Totally! This with DwarfStar delivers usable local AI (I hope!)
dannyw 11 hours ago|||
Usable local AI has been here for a while, esp on say a 5090.

You can’t treat Qwen3.6 like its fable, but if you prompt precisely and specifically it’s a great executor.

I actually found it refreshing to use more of my brain for once, and actually have to think deeper about what I’m trying to do, and how to build it.

genxy 3 hours ago||||
Parent is referring to https://github.com/antirez/ds4
Tepix 9 hours ago|||
What are your goalposts? Depending on your requirements, there have been many moments of usable local AI. More recent ones were gpt-oss 120b and Qwen 3.6 27b.
KaseyKim 10 hours ago||
hope that deepseek become better
ilaksh 9 hours ago||
Have you tried the one that was just released?
f311a 12 hours ago||
I've been driving flash model for 90% of my tasks. It's better than pro (for unknown reasons), very cheap and fast.

I try to keep changes under 1000 lines and drive architectural decisions myself, barely notice any difference compared to frontier models. The rest 10% is to spot bugs, security problems and to investigate better architecture, which flash can also do pretty well, I just cross check it.

Faster iterations are way better for me, I hate waiting for 5-10 minutes on small changes. I tried to use recent versions of Kimi and GLM, but they use too much thinking for no reason and are pretty slow because of it. I also often feed a lot of data to it, without worrying about hitting the limits: dependencies (to find bottlenecks in them), logs, performance dumps and so on.

Also, it will never complain about security guards, I've been using it to reverse engineer binaries.

embedding-shape 12 hours ago||
> Also, it will never complain about security guards, I've been using it to reverse engineer binaries.

Maybe I'm using too weak language in my prompts, but none of the OpenAI models I've used via codex has refused to reverse engineer binaries, is it supposed to? I'm sitting right now reverse-engineering a 3rd party firmware together with Codex and haven't hit a single guardrail. Meanwhile, I see people complaining about it rejecting non-security related prompts, are things so individual on the platforms right now or what's going on?

markasoftware 11 hours ago|||
Have you completed the identity verification? It's much more lenient once you have
embedding-shape 11 hours ago|||
Oh yeah, back in the GPT3.5 days I think, that's probably it. Thanks for sharing your hunch :)
flexagoon 11 hours ago|||
Weird, I'm using 5.6 Sol through Chinese resellers and it reverse engineers stuff just fine
johntarter 1 hour ago|||
That's hilarious that Anthropic and OpenAI can't even secure their products from Chinese resellers, how are they supposed to secure their models from being distilled?
genxy 3 hours ago||||
Because they have gone through identify verification, they are more likely to have done that to increase reputation score.
OkGoDoIt 2 hours ago||||
I think I’m missing something, what do you mean Chinese resellers? As I understand it, it’s difficult to even access OpenAI in China, how could they be reselling it? Do you mean something like openrouter but a Chinese version or something?
markasoftware 3 minutes ago|||
They usually resell codex subscriptions as api so it's cheaper than the official api
Scoundreller 1 hour ago|||
Proxies/vpns to enter/exit comms through non-banned countries.
dotancohen 2 hours ago||||
Do you have any suggestions for resellers? My Gmail username is the same as my HN username if you're not comfortable posting that here.

Thanks.

markasoftware 1 minute ago||
Here's a big list: https://vectoral.com/blog/token-relay-market
ngl999 11 hours ago|||
Probably not real 5.6 Sol
flexagoon 9 hours ago||
How so? The code and reasoning patterns certainly match, so does the intelligence level. I probably wouldn't know if they're secretly serving Luna instead of Sol, but it's definitely an OpenAI model.
realusername 11 hours ago|||
I got an account warning on OpenAI (waved after I complained) just because I was asking it how to root some >10 years old Android device.
dotancohen 2 hours ago||
Sonnet recently even suggested I root a three year old device and provided instructions. I suppose that is depends on use case. My use case was exporting data from an abandoned application running on an S24 Ultra.

The idea was that I could continue to use the application in the future and export the new data, not that I would be able to recover the extent data already in there.

mark_l_watson 9 hours ago|||
This is my experience also: DeepSeek v4 flash is good enough for most of my work and I like the fast response times. I buy tokens mostly from FireWorks.ai in the US, but I also prepaid for a large chunk of tokens directoy with DeepSeek.

I use OpenCode mostly (uses fewer tokens than Claude Code) and I am looking forward to the release of DeepSeek’s own coding harness.

regularfry 12 hours ago||
It's good but (at least on openrouter) it's got an annoyingly tight output token limit. So if it does get stuck in a reasoning pit, it won't work its way out of it in time.

It's replaced the Kimi models for me though.

u8080 11 hours ago||
Try to use DS platform directly - cheaper and better than openrouter, no subscription
dghlsakjg 5 hours ago||
Openrouter has cheaper inference than deepseek for the prior version of v4 flash, which I suspect will happen within a day or two with this version, and they don’t offer subscriptions. Are you confusing it with opencode?
u8080 3 hours ago||
No, maybe I've checked too much time ago, but direct DS was cheaper for all versions - also I had issues with openrouter providers availability
dghlsakjg 2 hours ago||
There are numerous providers offering full weight versions for a significant discount.

Availability and rate limiting can be an issue, though. I’ve found that constraining the providers works well to solve that though.

lionkor 12 hours ago||
I use deepseek for a lot of my personal day-to-day agent needs, and I will simply put this here and let this speak for itself, last 30 days:

- Cost: $4.55USD

- API requests: 3,467

- Tokens: 323,183,886

And as an engineer who leads a small team, I have very high standards for quality, and these carry across to my personal projects where I use deepseek. It has not disappointed at all for coding or review tasks. For everything else, use another model.

troglodytetrain 1 hour ago||
DeepSeek is amazing, they are, from a cost/benefit literally an order of magnitude or more better than the 'SOTA' models, and yet no one really talks about them.

I'm using them for my micro-saas, and they have made my niche economically profitable where as SOTA models are only slightly better for massively increased expense. Its truly impressive.

Word of advice to anyone, not all your use of LLM tech needs to be code/dev work related.

We are entering 'Web 4.0 era' or whatever you want to call it. Massive transformations of nearly every single business will and are being developed as the cost of intelligence as a commodity is falling through the floor...

a20eac1d 9 hours ago|||
Can you give more info on how you use/prompt those LLMs for code review and what kind of prompts you use?

I've had worse experiences doing it because the quality of answer has been quite bad, and I'm wondering if my methods are the reason.

lionkor 9 hours ago||
Yes, gladly! I have not yet open-sourced my skills etc., but I can give some insight and share a couple here.

Review is a skill, as in, a SKILL.md with a folder full of references:

- SKILL.md: https://gist.github.com/lionkor/161525be858d1d75db4c13c0f093...

- references/output-contract.md: https://gist.github.com/lionkor/8c68e33becef7a21f8408c7dc119...

- references/review-lenses.md: https://gist.github.com/lionkor/0a8b080fe45306213efddf3ebb75...

- references/review-workflow.md: https://gist.github.com/lionkor/d2d374b133ceb7e3660bd530ee72...

- references/section-rules.md: https://gist.github.com/lionkor/8a9e503adc7fd3697410cf021f27...

I'm aware that almost all of this is prompt voodoo, and there's no guarantee for the review to find anything or everything, but making it a dedicated skill and thoroughly observing the output thinking, tool calls, and result, lets me adjust these over time and fill the weak spots with even more prompting.

I use this skill by simply telling the agent something like "Review the changes on the current branch against origin/main, take special care with backwards-incompatible changes to the public API" or something like that.

I use `pi` (pi.dev) with a subagents extension, so that I can ask the agent to invoke a subagent to do the review, on work that the agent did.

For models, I use the highest possible reasoning on whatever model I feel like makes sense, usually this is GPT-5.5 or deepseek flash/pro, depending on the confidentiality of the codebase, on the highest reasoning always (for reviews).

I've also had success with a review checklist, though it doesn't produce an easy to parse (for humans) output: https://gist.github.com/lionkor/054ac2cf241e0765eee2383f0dba...

This is why my review skill mandates a very strict output contract. I need the output to be very easy to parse, and the output contract I've specified there does that.

In general I let <whatever the latest model of OpenAI's ChatGPT is> author and review SKILL.md and similar large prompts, usually with a ruleset like this, which is a 1600 line research artifact from a long GPT 5.5 "Pro" research session on prompt engineering: https://gist.github.com/lionkor/71498794d0a7d72173fc58766f25...

Does the review catch all issues? Not at all. Does it catch, usually more than one, important issue, across large changesets? Absolutely, and that's the point! :)

Feel free to ask me any questions, I'm also happy to share more about my setup via email or add you or anyone else to my private repos with more of these.

rzerowan 3 hours ago|||
Which languages are the reviews most successful on in your work so far? Im assumining js/python , would a similar output be feasible on lower level c++ or C#/Java.
lionkor 2 hours ago||
No js or python, it's mostly C#, Rust, and C++. They're helpful/useful on all of those, I've also used the review skill on shell scripts, GitLab pipelines, Powershell scripts, C, and a couple other odd things. It's almost never useless to run it on something, it usually finds things, even if they're sometimes not very major.
kekebo 5 hours ago|||
Thanks for sharing.
throwa356262 10 hours ago|||
What harness are you using to achieve that level of token caching?
lionkor 9 hours ago|||
I use pi, and, like the sibling comment, the caching ratio is fantastic. I work on C#, Rust, C, C++, shell scripting, and other areas.
k__ 9 hours ago|||
I'm using pi and my caching is ~99%.
brcmthrowaway 4 hours ago|||
OpenRouter?
lionkor 2 hours ago||
No, platform.deepseek.com for me! The caching is super important, not sure how open router performs, I haven't tried it.
marcus_cemes 1 hour ago||
Works out of the box with OpenRouter for most models. Some providers are a bit flaky, but DeekSeek (provider) has been one of the most reliable for me, no problems hitting >99% CH.
dakolli 11 hours ago|||
[flagged]
lionkor 9 hours ago|||
Aside from what squidbeak already said, I have to say the quality varies a lot, and goes above the "good" threshold enough times that using these models is worth the time and effort. Compared to something like GPT-5.4 to GPT-5.6 Codex models, they cost 50-100x less (fifty to one hundred TIMES less), and can do most annoying or well-specified tasks just as well.

The main difference is that deepseek is bad at prose, and Codex models are much more eager to use tools provided to them (which is usually fine, they often make like 3 todos via tool calls for a simple one step task which is silly though), and deepseek in general benefits from good instructions more than GPT models maybe do.

Deepseek becomes much better if you give it tools for asking clarifying questions, doing self-review with subagents, and so on. The more tools it uses, the better the signal-to-noise ratio in its context, and the more consistent the output.

You must still review 100% of the code as if it's trying to sell you insurance.

mcbuilder 6 hours ago|||
I love deepseek's (Pro) writing. I feel it's more nuanced and natural.
ai_fry_ur_brain 9 hours ago|||
[dead]
squidbeak 10 hours ago||||
An experienced software engineer (read his profile) praises the value he's found in Deepseek, and gives some real data showing how affordable that value is.

Then you, dakolli - out of generosity and minute-to-minute devotion to enlightenment - sacrifice time from your busy day to sit down (though perhaps that's been painful lately?) or stand up with your phone - and offer a profound, deeply thought-out counterpoint in the following form (and I'll paraphrase):

"Nah mate, it's shit. All LLMs are shit."

CamperBob2 3 hours ago|||
Look at his history -- he's just here to shitpost. Flag and move on.
applicative 8 hours ago||||
How one is affected psychologically by LLMs 'hallucinating' presumably closely tracks how one is affected by pathological liars. Most of the things pathological liars say is true and if you can put up with it, one can work with them and learn from them. But some respond with extreme hatred and avoidance. This seems like one of the responses that is reasonable ... as is acceptance and instrumentalization of the individual.
_superposition_ 9 hours ago||||
Among countless other experienced engineers... The cognitive dissonance here is frightening.
dakolli 8 hours ago|||
[flagged]
shock 8 hours ago|||
> I really can't begin to describe how stupid I think you are for thinking that having that statement in your bio makes you a qualified engineer.

Well, I guess having "Two PhDs, Three masters degrees. Expert in everything.." in your bio makes you very smart!

lionkor 8 hours ago|||
I agree a bio is a bad proxy, and I'm not saying I'm amazing but I'm not THAT incompetent. Most of my GitHub is pre-LLM era; github.com/lionkor

Edit: And like a lot of tools, HOW you use them is just about as important as the quality of the tool itself. Of course the tools can produce massive amounts of bad quality slop, they can also produce fast, focused edits that make sense.

dakolli 7 hours ago||
I'm not even trying to call you out, I'm sure you're solid.
lionkor 7 hours ago||
Oh I know, I'm just trying to say that I agree fully that a bio is a terrible proxy, while also showing what I think is a good proxy for skill (open source work before LLMs became this "good", for example). I'm a big proponent of show, don't tell, and LLMs have kind of ruined this.
ncphillips 11 hours ago|||
For doing what?
dakolli 10 hours ago||
everything.
seanmcdirmid 10 hours ago|||
Maybe it’s just your ability to use the tool and other people more proficient in using the tool are actually getting reasonable results from it.
ai_fry_ur_brain 9 hours ago||
[dead]
hlynurd 10 hours ago|||
There's a small relief in knowing that these junk comments are at least not AI generated (in all likelihood.)
innis226 11 hours ago|||
But doesn't it hallucinate a lot? Does that affect your workflow?
lionkor 9 hours ago||
It hallucinates plenty, about the same as Codex models and all other LLMs! I review all code it writes, thoroughly, check the test coverage, write tests myself, have other models/chats cross-check the work with a review skill (which I've shared in another comment in this thread).

Usually unreviewed code only gets committed if I really don't care, like for one-off scripts, which I sandbox with github.com/lionkor/sbh or run as an unprivileged user.

maweaver 3 hours ago||
You are using DeepSeek's services directly? Doesn't that end up sending at least snippets/chunks of code to a server where it is subject to Chinese government data access laws? Even if I was okay with that, my organization would never be. And even if they were, our partners/vendors/customers would not be. I think that's the sticking point for a lot of people.
lionkor 2 hours ago|||
Yes all of it goes straight to China! All of it is also open source or source available, or will be in the future. The stuff that isn't is trivial enough.

For any work with protected intellectual property, I use other providers, for the contractual guarantees, but I think it would be silly to think that OpenAI or Anthropic are not training on literally all data they get. How could you ever tell if they did? They can just claim the data was mislabelled, or ignore the accusations. If you have serious IP, use only local models.

anigbrowl 2 hours ago|||
So go with some other provider who hosts the models like OpenRouter if are worried about this.
kmarc 12 hours ago||
Essentially I'm running everything on flash now inside pi. With the correct set of MCP servers, context reducer tooling and skills it can implement any task I throw at it. Some sessions take 30+ turns, but it's fast and cheap; all this in an hour, with ~$0.5 cost.

(TBH though, in my multi-subagent workflow I do use other, more expensive models for planning, reviewing, oracle-ing)

I haven't used our slow opus subscription for weeks.

(Also set up an OpenWebUi self-hosted chat that works from my phone, has some mcp and skills. fully replaced perplexity. Monthly cost ~$18 for hosting and subscriptions)

lionkor 12 hours ago||
I want to second this, I use the same setup (pi + deepseek, with lots of custom tools for tracking TODOs, doing things with less tokens, etc, and with a SOTA model for very difficult tasks) and it's all I need it to be.
peperunas 12 hours ago|||
Do you have any recommendations of such extensions for pi?
sdesol 3 hours ago|||
This is self promotional but I am working on making pi extremely enterprise ready with:

https://github.com/gitsense/pi-brains/tree/staging

The README is being worked on but the three videos should give you a good sense of what it can do. Pi is also what makes what I will demo in

https://github.com/gitsense/chat/tree/update-readme

possible. Since Pi exposes so much, it is very easy to build advanced tooling around it to help easily grok hundreds of tool calls to help you understand what they agent knows and what it has tried.

kmarc 12 hours ago||||
Recommendation? No. Just go with the passive-aggressive advice "let pi build it for you". :-)

To be more constructive, what I did (as an experiencd SWE but a complete noob to agentic coding): went to pi.dev's extension marketplace and looked into all the new shiny stuff. Subagents, mcps, context and memory optimizers, skills. Using the most popular ones (not necessarily the best ones)

It was like 15years ago learning the new mindset of vim (and spending a ton of time to customize it to my workflow). My understanding is that Claude and opencode doesn't give you this flexibility.

Learning all these stuff drove me to also set up openwebui, and it was such a successful private project that I implemented it at work (with jira/confluence/bazel query access) and management said "we need this by tomorrow".

I believe the models matter not that much anymore. The "harness" does. (unless you just want to vibe code. Thebn, throw crap at fable and call it a day)

javier123454321 7 hours ago|||
I found that opencode absolutely lets you do everything that pi does. With a few niceties in the UI on top.
epolanski 4 hours ago||
Opencode doesn't even allow providing your own system prompt.
peperunas 12 hours ago|||
Got it :-)
rurban 12 hours ago||||
Happy with oh-my-pi (omp)
peperunas 12 hours ago||
I looked into it but, at least for me, it goes against the entire philosophy of pi :-)
WithinReason 11 hours ago||
the entire philosophy of pi is to be extensible
jeremyjh 9 hours ago||
omp is not really pi. Its a very distant fork of it. And its philosophy is completely different, it includes everything you need except workflow skills. Personally, I think it is the best agent coding tool by far, and its what I settled on after trying 6+ different tools.
sergiotapia 5 hours ago||||
I've had a lot of fun with https://omp.sh/ - consider it like zsh -> ohmyzsh

Lots of sensible defaults and good tweaks/settings you just don't worry about.

try-working 12 hours ago|||
pi-role-model
kzrdude 11 hours ago|||
Do you notice any improvements with this update?
kmarc 10 hours ago||
Not sure, haven't tried it today.

Last night it single handedly implemented a feature after a grilling session, and came back with the red-yellow-green risk assessment points that I mostly saw with anthropic models. I had to check if I'm using the right model, but it was DS4Flash.

So maybe I was using it already?

the_lucifer 8 hours ago||
Possibly depending on who's routing it, since Openrouter seems to have added a separate slug for this, and they don't have an alias router for the DS models yet, so it shouldn't have swapped over to it in case you use them for routing.

If you are directly using it through Deepseek, they do mention that the flash slug has migrated over

rdsubhas 11 hours ago||
How can one set topp and temperature in pi?
embedding-shape 11 hours ago||
You build it as an extension or create your own fork/copy of pi and add it. This basically goes for most things in pi, except the most basic stuff. It's basically meant for people who think "I'll just add that myself" rather than expecting it to be there out of the box or finding other's solution to it, for better or worse.
wkcheng 13 hours ago||
If the benchmarks are real and reflect actual use, then this is an insane model. This 300B model outperforms the previous DS4 Pro preview model (1.8T params), and it looks like it outperforms GPT 5.6 Luna too. And it's still cheaper than Luna, even with the price decrease.

Crazy.

gr_norm 2 hours ago||
OpenAI must've known this was coming, hence the Luna price drop. This competition is amazing!
cbg0 11 hours ago||
On DeepSWE Deepseek is 54.4% and Luna is 67%
benjiro29 9 hours ago||
But on Terminal bench, its

* DS4 Flash: 82.7

* GPT 5.6 Luna: 75.7

For reference, that puts it on the third spot behind GPT 5.5 and Fable 5. For some reason GPT 5.6 Sol is not showing in the leaderboard. If it did, then DS4 Flash was number four.

The thing is, even if Luna is better in DeepSWE and has the 80% discount. DeepSeek is still cheaper.

--------------

DeepSeek V4 | Flash GPT-5.6 Luna (New)

--------------

Input (Cache Hit) $0.0028 $0.02

Input (Cache Miss) $0.14 $0.20

Output $0.28 $1.20

--------------

Both Luna and Flash are heavy on the reasoning > output. And the cache hitrate + prices also matter.

Reality is, you can not go wrong with Luna or Flash at those prices. And remember, DeepSeek V4 Pro is still in the rafters, what is ironically closer to Luna's new price.

flashblaze 8 hours ago||
Agreed. The only thing missing is multimodal capability
dannyw 11 hours ago||
Deepseek and moonshot are the only two providers I consent to training for.
arizen 10 hours ago||
Why not also Qwen?
dannyw 10 hours ago||
Qwen/Alibaba have stopped doing open weights releases for a while. No grudge or anything, I'm certainly not going to look at a gift horse in the mouth, but both DeepSeek and Moonshot have been very consistent with open weights as well as sharing actually detailed research.

In terms of open research, China has absolutely overtaken the US.

anon373839 10 hours ago||
Qwen has a pinned tweet stating that 3.8 will be released as open weights soon. I guess it remains to be seen, though, if they’ll do the smaller model sizes or only the big 2.4T one.
SXX 9 hours ago||
Most interesting part IMO is whatever Qwen open weight will include image / video capabilities since its where they are strongest.
mcbuilder 8 hours ago||
Also base models versions are very useful for researchers, since they allow for a cold start to post training.
siva7 6 hours ago|||
CCP will be happy! Go on and share all your data with them..
neya 4 hours ago|||
Yeah, because, your data is best collected by the US labs, right?

US or American, don't trust anyone, these are open weight models. Host them yourself if you feel strongly about privacy (as you should). I honestly don't know of any OSS open weight models from the US labs as good as Kimi K3 or Deepseek V4 though.

fdsjgfklsfd 5 hours ago||||
Yes, they trained on my open source code and Wikipedia edits and Stack Overflow answers that I shared freely, so it's only fair that they release their models as open source and share back to the community that created them. I only allow my training data go to open source models.
gr_norm 2 hours ago||
Yeah, at least when I share my training data with labs releasing their models openly (Chinese or American or otherwise) it's nominally so an even better open model will land in my hands in the future.
chorizo 5 hours ago||||
I’m developing open source tools, so happy to have all my context traces be public if it helps improve future models.
anigbrowl 2 hours ago||||
I will! I pay their sales too and consider it quite reasonable.
genxy 3 hours ago||||
At least you get based models back at some point. You supply data, they supply compute and open weights.
bel8 2 hours ago|||
Yes I will tyvm.

They release the models back for free.

applicative 8 hours ago||
Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.
satvikpendem 6 hours ago|||
He explicitly said he wants open weight models at his recent speech at an AI conference in Shanghai: https://news.ycombinator.com/item?id=48970449#48970784
applicative 2 hours ago||
The Chinese equivalent of the expression 'open weights' appears nowhere in this talk. You believed the Western press. Chinese AIs are not typically open. The most important by far is Bytedance.
Petersipoi 35 minutes ago||
The second China shuts down open weights is the second they lose the AI race. Nobody is going to use Chinese models if they're closed, besides hobbyists who don't have anything they care about getting stolen.
gravypod 8 hours ago||||
What would the benefit of this be? If China stops open weights, US labs still have the intelligence frontier. Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't.

Undermining out entire economy by subsidizing the release of DIY versions of our main economic drive sounds like a huge win for China.

applicative 6 hours ago|||
> Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't.

Yes this is why I referred to 'months' and was downvoted by people who can't distinguish their politics from reality. You are restating exactly the text you are criticizing.

This is the nature of mechanical parrotlike repetition of propaganda:

a) You can have 'frontier' models with closed weights. Your sentence is basically a contraction within itself, again, the weight of ideology.

b) No one actually knows whether or not the closed weights of Bytedance are beyond all existing frontiers. Again the weight of ideology blinds you to the fact that the real 800 lb gorilla of Chinese Ai is more closed than Anthropic.

ricardobeat 6 hours ago|||
The latest ByteDance Seed model is behind Opus 4.6 in performance. It’s not even in the conversation atm.

The parent made a totally coherent commment that China closing off their open weight models would cede the frontier (and majority of AI users worldwide) to the USA. Nothing you said counters that point.

applicative 2 hours ago||
Bytedance models completely pervade business use across China. It is the OpenAI of China, whether you think its a good model or not.
gravypod 4 hours ago|||
If they did have a super powerful secret AI why would avoid:

1. Selling access to the US? 2. Ship more and better software? 3. Talk about it publicly?

If you had a secret AI better than anything currently available the mere mention of this would crater the US tech investment sentiment.

They could even host these Chinese models through AWS and sell at-cost inference on Bedrock and obliterate the US model companies.

anon373839 8 hours ago||||
Reuters was reporting that rumor. And then Xi made a public appearance at a conference in Shanghai where he said the opposite of that rumor.
applicative 7 hours ago||
No in that very speech Xi stated explicitly what amounted to: 'of course when we get a Mythos, it will be a state secret'.

In fact the overwhelming weight of AI use in China, the chatgpt so to say, is Bytedance's AI which is absolutely closed and uniquely opaque.

The press treatment of these matters was no good and they are slowly walking it back, e.g. NYT yesterday finally actually read the speech.

anon373839 6 hours ago||
Do you have a source for that? The NYT article does not quote Xi making that statement, and I can't find a source for that. In fact, Xi delivered veiled criticism of the security posturing the US has done:

> We should put in place laws and regulations, technological monitoring, early warning and emergency response systems in order to strengthen the line of security, prevent abuses and malicious use, and ensure that AI is always under human control. In the meantime, we should jointly oppose overstretching the national security concept in the field of AI and placing one country’s security over that of others.

There was a blog post by a state-linked broadcaster cited in the NYT piece:

> “China supports openness, but this does not mean it advocates for the unconditional proliferation of all capabilities,” the blog said.

The article also quoted an American journalist:

> “If these models do reach those dangerous capabilities, they are not going to let it be a free-for-all in terms of releasing them," Mr. Sheehan said.

But a distinction has to be drawn between real dangers (which many people believe LLMs have not actually shown, to date) versus "dangers" hyped up marketing purposes or domestic regulatory-capture motives. Presumably the Chinese government is less interested in the latter.

applicative 6 hours ago||
"Dangers" may indeed be hyped up for marketing purposes or domestic regulatory-capture motives. But as Xi stated, they will become real.

It is just a question of being unaffected by the motives of the speakers, which is what adults learn to do.

brazukadev 6 hours ago|||
replying the request for sources with "as xi stated" without bringing the sources again is quite misleading. It seems you have nothing to back what you are saying.
applicative 2 hours ago||
"google is your friend"

The person I am replying to actually believes that Xi actually spoke of "open weights", for example. This is the AI information space we live in.

anon373839 6 hours ago|||
> But as Xi stated, they will become real.

But the speech doesn't say that? I'm looking at the full text. For example:

> We should take seriously the various types of inherent and secondary risks that AI may trigger.

The risk portion of the speech is focused more on application-level risks than model capabilities. Which is a much more sensible regulatory framing than what Silicon Valley has proposed.

applicative 2 hours ago||
The claim that the speech was somehow supporting open weights was pure Western imposition. China does not support open weights generally, and Xi makes plain that some systems of weights - perhaps still future - would have inherent risks, same as Amodei says of Mythos, whether or not he is bull_ing about it.
GTP 7 hours ago||||
We can only wait to see if this is true. Another possibility could be that it was an answer to USA's government considering a ban on chinese models. In this way, he fueled the discussion around the importance of open weight models.
applicative 7 hours ago||
Xi actually said that models of the strength to pose security issues will be permanent state secrets -- in the same speech that the first western takes translated as actually using the words 'open weights' which of course nowhere appeared. Now: when will "models of the strength to pose security issues" appear? Maybe my inference 'months' is wrong, but they seem to be advancing rapidly.

The administration blather about banning open weights is characteristically confused. Xi has already stated (what is obvious) that he will ban security-endangering weights and keep them a state secret.

darkwater 8 hours ago||||
What makes you state this?
riskd 7 hours ago|||
Decades of xenophobic propaganda
applicative 7 hours ago|||
No, reading Xi's speech. They aren't going to give their Mythos - presumably a few months away - to the Sinaloa cartel or Uighur hackers or the US military or etc etc
h8hawk 6 hours ago|||
They have already released their models (GLM 5.2 and K3). This is why Amodei and the rest of the effective altruist lobbies are trying to pressure the US government to both ban open-source models and pressure China to do the same.

I’m curious to see why you’re so eager to ban open-source LLMs while being okay with Anthropic’s control. Do you think they’ll rebel against you?

sudosysgen 3 hours ago|||
Mythos is Fable without safeguards. Open weight models are already essentially at the Mythos level, and with extra post training a well resourced actor could deploy would almost certainly be significantly better than Mythos at hacking. Whatever their threshold is, it's much higher than that.
1234letshaveatw 7 hours ago|||
Wow! Amazing observation! Wait, do you think there could have been any pro Chinese propaganda? Perhaps even some taking place at the present time? I think it could be possible?
Zetaphor 6 hours ago||
I'm sure there's plenty, but if you put on your rational hat and think about the motivations of the Chinese government you'll realize this claim makes no sense, and there's plenty of evidence and good reason saying the opposite
1234letshaveatw 5 hours ago||
you are probably right! nobody does more domestic discourse shaping than China, but good reason says none of that would take place externally
baublet 8 hours ago|||
China bad
applicative 7 hours ago||
I'm a) quoting Xi's speech and b) projecting rapid China development to Mythos level. Either you think China will never get to 'Mythos level' because you think China bad; or you think they will release the weights, because you think China bad. Its incredible the weight of ideology over facts in this discourse
satvikpendem 5 hours ago||
You haven't quoted anything actually. Provide sources to your claims.
codemk8 2 hours ago||||
Me.
1234letshaveatw 7 hours ago||||
That isn't the Chinese way. They are much more focused on undermine and extinguish. Just look at the European car industry- on its way to being non-existent after the market was flooded, bye bye manufacturing base. Undercut the US AI providers and wait them out, they will go private after they have a stranglehold
bwfan123 5 hours ago|||
> Undercut the US AI providers and wait them out

You could have said the same of linux - that it extinguished proprietary OSs at least for server usecases.

1234letshaveatw 2 hours ago||
true, but the motivations of OSS contributors are benign (overall)
applicative 7 hours ago|||
Even Xi's speech was as paranoid as Dario Amodei about the possibility of Chinese AI achieving something at the level of say Mythos. If you seriously believe such a thing will be on Hugging Face I really don't know what to say.
fdsjgfklsfd 5 hours ago|||
DeepSeek-V4-Flash-0731 scores higher than Fable 5 on Terminal-Bench, and Fable 5 is Mythos, correct?
criley2 6 hours ago|||
There is already something on HuggingFace at the level of Mythos. It's called Kimi K3 and it's running laps around Opus5, Fable5, and Sol56 at cybersecurity.

It's so good that the US government is rushing to ban all Chinese models as fast as they can.

applicative 6 hours ago||
Fine, you believe Anthropic was lying about Mythos. In fact no view of the matter is needed to formulate the proposition above. It is clear China will soon have a model with the powers imputed to Mythos but which you deny of actual Mythos. It is a trial to deal with this level of ideological blindness; 'ideology has no outside'. If someone says 'when they get Mythos...', declare the possibility of Mythos to be the real lie. But it is quite possible and will be coming in months.
zozbot234 6 hours ago|||
Reading between the lines, it's fairly clear that Mythos Preview only got its reported results in offensive cyber thanks to a highly specialized harness and humongous amounts of test-time compute. We also don't know what kind of post-training/fine-tuning it got on cyber, Anthropic only stated that it wasn't specifically trained for cyber-offense but that leaves a lot of scope for training on other related things. Overall, the Mythos story is a lot less interesting than most people might like to claim - but note that this doesn't require any actual lying.

The interest around K3 is a lot more defensible because we actually know what the architecture looks like, how much effort it takes to get it to run, and what kinds of results it gets on cyber evaluations. And no, it's nowhere near "dangerous" enough to where people might honestly want to ban it for real safety reasons. It does a good enough job at fixing cyber issues, but that's hardly a safety concern.

potwinkle 5 hours ago|||
[dead]
ReptileMan 7 hours ago|||
It is true. And only Chinese Han people can continue work on what is already released. But they will be forbidden. /s

Sarcasm aside - if the community can't get their shit together to continue improving what is currently public, well - we don't deserve free stuff and open weights.

ggcr 13 hours ago||
Woah, a 200B model competing with GLM-5.2 and getting close to Opus 4.8. Quite impressive.

If those numbers translate well to its general capabilities, with the great caching DeepSeek has, I feel like this model will get tons of usage.

lostmsu 9 hours ago||
Not just 200B model, it is only 160GiB.
ignoramous 8 hours ago||
MiniMax's another lab that's known for relatively smaller models (their latest, M3 is 295b) that punch way above its weight.
heyalexej 1 hour ago||
Long story, I have humongous zai GLM 5.2 token budget that I'm using in a similar fashion as many comments explain here. GPT 5.6 or Fable 5 for planning, GLM for implementing, researching, extracting and many other tasks I consider grunt work. Very happy with the performance, speed isn't all that good though. I'd be curious to hear from someone who works with both, DeepSeek and GLM side by side.
Goranek 12 hours ago||
Kimi K3 (instead of Opus) for expensive stuff, DSV4 Flash for tasks (instead of Sonnet)?

Does this make sense?

baalimago 12 hours ago||
Looks like they will release an updated version of deepseek-v4-pro soon, which most likely will beat kimi k3 at both intelligence and cost (judging by the vast improvements to dsv4 flash)
geek_at 12 hours ago|||
it does! I'm always amazed when I use DSV4 flash for coding or server checks and after an hour of working with it my (pure api call) balance is about 30 cents
try-working 12 hours ago||
[flagged]
thirtygeo 13 hours ago|
For both US and China models - what standard security checks and QaQc are you all doing? We're running small gamuts to test for unsolicited jailbreaks (model jailbreaks you) and incorrect records (Fake Accuracy - as Easter Egg or common thread) meaning falsified logic or information cooked in by the developers, rather than the training data speaking for itself
More comments...