Top
Best
New

Posted by snehesht 11 hours ago

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s(github.com)
550 points | 268 commentspage 4
paulez 7 hours ago|
Pretty impressive so far, but needs more testing.

It is more useful than Qwen3.8:27b (which is already quite good) and runs faster on my 7900 XTX / 64 GB DDR4 system.

Local LLM is getting more exciting every day!

coderbants 4 hours ago|
Interested to know throughput on 7900 XTX and what setup you're using?
ai_ja_nai 9 hours ago||
I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?
kennywinker 5 hours ago||
Yeah, -10% accuracy (probably more like -20% in reality) sucks, but only if you could be running it at 100%.

That's the exciting part of this - before the best you could run on <24gb vram was qwen3.8-27b at q4 quantization. Now you can run a nerfed 125B parameter model on under $800 of hardware, and it beats a less-nerfed 27b model.

ai_ja_nai 9 hours ago||
(64GB not VRAM, I meant) I also see people claiming fast performance on a 128GB machine, which is not exactly consumer hardware)
hypfer 10 hours ago||
Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?

The Readme doesn't say, but it's all AI generated, so..

eliaskg 4 hours ago|
I wanted to know as well. KV cache is Q8 by default but can be set up at full precision in the config.
hemedanmert 6 hours ago||
I need a version of this that runs 3.8 27B on 8 gigs of VRAM

amazing project, congrats on the launch

esafak 10 hours ago||
Has anyone calculated the effective intelligence of these quantized models?

I think publishing benchmarks with quantized models should become standard practice.

nsagent 10 hours ago||
See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

  We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

[1]: https://arxiv.org/abs/2608.08188

merbanan 10 hours ago||
I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.
mkl 10 hours ago||
There's some info in the README, including:

> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

https://github.com/Niko1221/Strata#which-model-should-i-pick

nicce 10 hours ago|||
I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?
kennywinker 5 hours ago||
125b at q2 is ~80gb

27b at q4 is ~16gb

So from a raw amount of data, qwen3.8-flash-next wins easily. But flash-next is an MoE model, so it only has 6b parameters active per token, vs 27b's dense 27b per token. So 27b@q4 uses ~16gb of weights per token, and flash-next uses about 4gb of weights (125/80 * 6).

But those numbers don't really tell us anything useful, because there is an interplay between total model size and active parameters and intelligence that isn't obvious or simple.

(sizes are based on the unsloth quants, not the coder variant, but the idea holds - this isn't calculatable with simple math, you gotta test them and see)

javier2 10 hours ago||||
ok that is getting interesting!
nisarg2 10 hours ago|||
92% is halfway to 99%

Holds up pretty well

jameslholcombe 4 hours ago||
I might try combining this a FreeToken
b212 9 hours ago||
I tried to run Qwen 3.6 27b locally a few months ago and all those synthetic tests do tell you something and quite a lot of people were very excited about that model but honestly? It wasn’t even close to default mode in Cursor or Sonnet at the time.

I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?

rpdillon 5 hours ago||
> Qwen 3.6 27b locally a few months ago

I've been doing local inference for a couple of years on the side, and I'm astonished at the number of variables you need to have control over to get a reliable result. Inference engine, model parameters (top_k, temp, MTP-enabled/not), quant level, and harness all have a big impact on the results.

DS4-0731 at 2bit on llama.cpp (ROCm) and 250k context with omp.sh has been consistently reliable for me, just a bit slow (10 t/s) compared to what I'd prefer. Trying out DwarfStar today (benching it right now) to see if I can get better speed, but otherwise I've found it to be great on my side projects that are smaller (up to 10ksloc).

There could also be a domain issue - I tend to do lots of web programming and sysadmin work in these projects; if your work is more esoteric, it might not be nearly as good. I haven't tested much outside of my narrow domain.

ApatheticCosmos 9 hours ago||
I started using Claude right before 4.5 came out, and 4.6 is where it turned a corner for my use.

Qwen 3.8 Flash Next is there. 3.8 27b is fairly close.

I'm excited to see what Qwen 4 will bring.

I'm running on a 128GB Strix Halo for Flash Next and an Intel Arc Pro B70 (32GB) for 27b.

MattyRad 7 hours ago||
I just used 5.5 xhigh reasoning to make a massive implementation spec (for a vibey throwaway project/exploration, not anything important, burned 80% of the 5h window), now my Strix is in the process of implementing it.

I think there's probably low-hanging fruit to outsource reasoning from rote read/writes.... Just speculation though, I'm not a token optimization expert.

Neywiny 9 hours ago||
I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.
halJordan 9 hours ago|
100% not an llm problem. Llama.cpp, which only recently started taking large amounts of ai code has had this problem for years.
Jeeetendra 9 hours ago||
getting it to fit is impressive, but i'd want to compare the smaller quants on a real coding task before picking one. how much quality do you lose going from IQ3_S to Q2_0?
Luker88 8 hours ago|
I have fairly limited HW, so i tried standard llama-cpp and qwen3.8-flash-next, unsloth quants.

Q1 was producing some garbage at times, generating wrong urls on webfetch, then convinced itself there was some url rewrite in the middle. With IQ2 it happened much less but still happened, and once it would all webfetches became like that. IQ3_XXS is the maximum I can run: I don't have problems anymore, though I have less available context window.

Jeeetendra 8 hours ago||
that's a useful comparison - i'd take less context over broken tool calls, though it'd be interesting to see if IQ3_XXS holds up on longer coding tasks too.
bt1a 7 hours ago|
80 t/s w/ 3090s and 3.05bpw exllamav3
More comments...