Top
Best
New

Posted by snehesht 10 hours ago

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s(github.com)
532 points | 262 commentspage 3
pilooch 8 hours ago|
My goto private setup, runs ~50t/sex on a dgx spark with sglang, nvfp4. Excellent model.
apitman 7 hours ago|
This is a meaningless metric without knowing how long the sex takes
swiftcoder 7 hours ago||
30 seconds at most
zkmon 8 hours ago||
>> The model is a team of 24,576 small specialists ("experts")

That's a neat number (576 is the square of 24). Ofcourse it must have come from 24 * 2^10.

kzrdude 1 hour ago|
I have to say that "team of specialist" and "model of experts" as explanations and names give a quite misleading picture about how it works, at least I thought so when I learned about how it worked.
gdevenyi 10 hours ago||
I had this working with the FreeToken inference engine a month ago when they launched.

https://github.com/FlashML-org/FreeToken

ryan_glass 9 hours ago||
Anyone know how it compares to GLM 5.3 for real world use?
mapontosevenths 8 hours ago||
À 2 bit quant will (at best) get you about 80% of the full models memories. That's from a purely information theoretical sense. IRL it's worse than that.

Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.

So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.

FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.

One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.

alienbaby 8 hours ago||
Terribly
mark_l_watson 8 hours ago||
Qwen 3.8 Flash Next is amazing. I only have a 64G Mac so I have to run Sushi project’s 3 bit quant. Amazing results with pi-dev. More for fun than anything else, but I am trying to do as much as possible with local models, now rarely falling back to a paid deepseek-4.1-flash API.

Progress on running local models has been amazing.

fsiefken 7 hours ago||
Yes, I am running the same on a 64G mc. It's good, but slow at 25 tps on average! I want > 100 tps - but I don't have $5k to spare for an m5 ultra or an nvidia setup.

So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2

https://huggingface.co/nathansutton/Qwen3.8-27B-Ternary-Bons...

or a MoE retrofit like Qwen3.8-35B-A3B with or without mtp

https://huggingface.co/NovaeonStudio/Qwen3.8-35B-A3B-Distill...

https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...

generalizations 8 hours ago||
I haven't seen much in the way of benchmarks of those smaller quants. How does it compare to e.g. various generations of Opus?
sweetboy 5 hours ago||
I think it would be great if you could try models with lesser parameters that could fit on 6GB VRAM-ish, which could work for "gaming laptops" as well.
kennywinker 5 hours ago|
There are options. If you have fast CPU RAM (ddr5) and PCIe bus you can run Qwen3.6-35b-a3b at good speeds+~100k context (I have a friend who is running this setup). If not, you're stuck with the much smaller models: Ling-3.0-tiny, Spark-X2.5-4B, LFM2.5 in 8b-a1b or 2.6b, and FrogNano-4B-2609 looks promising. But 6GB is a tough squeeze, 8gb is a lot cleaner, and if you have 12gb you can run strata like OP.
paulez 7 hours ago||
Pretty impressive so far, but needs more testing.

It is more useful than Qwen3.8:27b (which is already quite good) and runs faster on my 7900 XTX / 64 GB DDR4 system.

Local LLM is getting more exciting every day!

coderbants 3 hours ago|
Interested to know throughput on 7900 XTX and what setup you're using?
hemedanmert 5 hours ago||
I need a version of this that runs 3.8 27B on 8 gigs of VRAM

amazing project, congrats on the launch

ai_ja_nai 8 hours ago||
I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?
kennywinker 5 hours ago||
Yeah, -10% accuracy (probably more like -20% in reality) sucks, but only if you could be running it at 100%.

That's the exciting part of this - before the best you could run on <24gb vram was qwen3.8-27b at q4 quantization. Now you can run a nerfed 125B parameter model on under $800 of hardware, and it beats a less-nerfed 27b model.

ai_ja_nai 8 hours ago||
(64GB not VRAM, I meant) I also see people claiming fast performance on a 128GB machine, which is not exactly consumer hardware)
hypfer 9 hours ago|
Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?

The Readme doesn't say, but it's all AI generated, so..

eliaskg 3 hours ago|
I wanted to know as well. KV cache is Q8 by default but can be set up at full precision in the config.
More comments...