Top
Best
New

Posted by snehesht 9 hours ago

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s(github.com)
514 points | 255 commentspage 2
Luker88 8 hours ago|
Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.

Surprisingly useful as long as you can leave it running a couple of hours at the very least.

While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.

londons_explore 8 hours ago||
Remember that a hosted AI company only needs ~1000 bytes per context token per user (ie. 100mb per user for typical coding - VRAM during inference, and moved to regular RAM or SSD whilst running a tool call)

Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.

eurekin 7 hours ago||
> ~1000 bytes per context token per user

Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb

jburgess777 7 hours ago||
The smallest I have seen is DeepSeek 4.1 flash at 890 bytes per token.
bitexploder 6 hours ago||
Highly recommend the RCO-GSQ quant by ITSA btw. At IQ3_XSS it is within one point of the fully unquantized model.
Luker88 5 hours ago||
will try, thank you for the pointer!
prettyblocks 9 hours ago||
I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).
zkmon 2 hours ago||
I don't get it. It's file size is about 6 times larger than 27B model for the same quant, but the performance improvement is hardly 10% across all benchmarks, according the metrics on it's hf page. Why should one devote so much more hardware for so little benefit?
nialv7 8 hours ago||
There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...
bitexploder 6 hours ago||
Not sure I have a strong opinion but I am sort of okay with current state of affairs. I am optimizing for V100. EOL cards on EOL CUDA. Llama is a good enough base for this. A couple weeks of grunting at Claude has gotten the inference /fast/ for my uses. 150-160 t/s on 2 GPU for 27B and 125 t/s on Flash Next. Asking them to upstream every random feature does not make sense. They sacrifice a lot of speed to maintain stability and a reasonable feature set that works across a diverse range of models and systems. They could maybe merge some features like this and gate them on flags a little faster, but you can cobble together what you need and the big models can figure out how to make it fast.
mkatx 39 minutes ago||
Does anything else support Pascal gpu's though?
Tepix 5 hours ago||
All headlines about LLM performance MUST have the quantization also mentioned in the headline.

You know, so you're not wasting your time like in this post.

rpdillon 4 hours ago|
Quants vary by model. DS4 is very credible at a 2-bit quant. Not sure about Qwen 3.8 Flash Next; I run it at a 4-bit quant and it's too slow, so I'm trying out DwarfStar today to see if that improves things.
hecturchi 5 hours ago||
- Tiny context size or hours to load it

- Hard to benefit from thinking and preserve thinking given token cost.

- Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.

- K/V quants probably quantized too make things less accurate.

Useful would be combinations with:

- Full context size so it can code and think a bit.

- Draft MTP <= 2 so it doesn't trip

- Q4 quants or better so its accurate

- q8 cache or better so it stays accurate.

- 20 token/s so it finishes while reviewing previous step.

- 1000 tokens/s context load so compactions don't waste 10+ minutes.

- And enough left RAM for 50+ context checkpoints so that it can progress quuckly.

Closest you have is Qwen3.6-35B-A3B-MTP.

Latest gens (Qwen3.8 and co.) are just too big for low specs. 27B dense models seem to be ok for integrated >=92 GiB RAM.

Source: I have low specs and tried them all for agentic use + coding.

kennywinker 4 hours ago||
Everyone's definition of usable is different, but I disagree with your estimation of the specs required to be useful. I am able to do useful coding on Qwen3.8-27b, with 100k context and 16gb VRAM, q8 cache. I feel I'm living right on the cusp... my GPU is old (2016, pascal), so to get usable speeds I have to drop to a Q2 quant - which still gets stuff done, but the difference with q4 is noticeable. Q3 is close enough I don't really notice the difference between it and Q4, but it's too slow on my system. More context would be nice, but it's not that hard to work within ~100k.
ohyes 4 hours ago|||
I’m using qwen3.8 quants (q3) effectively on a 5070 ti. A lot of it is about guardrails, but you also have to figure out how far (and in what ways) you can push a given model.
mirekrusin 4 hours ago|||
27b runs perfectly fine on 2x 24GB at ~100 t/s (4090) with speculative decoding on 8 bit quants
hecturchi 3 hours ago||
I bet! Just 2x24GB is not super basic hardware imho.
ranger_danger 5 hours ago|||
What about Bonsai 2? You can fit Qwen3.8 27B on an 8GB GPU with it, and upstream llama.cpp support is already being worked on (they just got System1 support too).
Aurornis 4 hours ago|||
The Bonsai models are really bad when you actually use them for more than short responses.

Their marketing made it look like a breakthrough, but in my experience it’s just the next step down from the Q2 quants in both size and quality.

Q2 quants are already not very useful in my experience. The Bonsai models are even worse.

If you only need 80% plausible outputs that don’t need to reference a lot of context they can be useful. If you try to use them for real tasks it feels like time warping back to 2023 when you LLMs were barely useful if you babysat every word of the output.

ranger_danger 3 hours ago||
Have you actually used Bonsai 2 though and not just the original Bonsai? The experience is vastly improved but still requires a custom llama.cpp fork to use as of right now.
arcanemachiner 4 hours ago|||
The ternary model? Hopefully those are worth a damn in a few years, but currently just an interesting toy from what I understand.
qeternity 3 hours ago|||
> Flash attention with who knows how much draft ensures it makes the same mistakes all the time and cannot call tools, format output or follow basic guidelines reliably.

> Draft MTP <= 2 so it doesn't trip

I am not sure you understand what either of these things do.

Do you think that FA or MTP are lossy?

hecturchi 3 hours ago||
I mixed FA and MTP wrong in my original post, thanks for pointing it out.

My experience is that draft-mtp=2 gives 25% improvement in tokens/s but the model is unable to call tools with the right arguments reliably. I have since then gone for smaller quants at draft-mtp=1 and that problem is at least gone so far.

An agent needs to repeat tool calls in the right way, so cannot penalize repetition. This however leads to the agent retrying the wrong calls constantly.

dang 3 hours ago||
Can you please make your substantive points without snark or swipes? This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.

There's a lot of good information here but the comment spoils itself by coming across as aggressive in this way.

This is particularly a problem when responding to someone else's work. We need commenters to point out problems respectfully, not put down what other people have been making.

hecturchi 3 hours ago||
Sorry, edited
dang 1 hour ago||
Appreciated!
b212 8 hours ago||
I tried to run Qwen 3.6 27b locally a few months ago and all those synthetic tests do tell you something and quite a lot of people were very excited about that model but honestly? It wasn’t even close to default mode in Cursor or Sonnet at the time.

I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?

rpdillon 4 hours ago||
> Qwen 3.6 27b locally a few months ago

I've been doing local inference for a couple of years on the side, and I'm astonished at the number of variables you need to have control over to get a reliable result. Inference engine, model parameters (top_k, temp, MTP-enabled/not), quant level, and harness all have a big impact on the results.

DS4-0731 at 2bit on llama.cpp (ROCm) and 250k context with omp.sh has been consistently reliable for me, just a bit slow (10 t/s) compared to what I'd prefer. Trying out DwarfStar today (benching it right now) to see if I can get better speed, but otherwise I've found it to be great on my side projects that are smaller (up to 10ksloc).

There could also be a domain issue - I tend to do lots of web programming and sysadmin work in these projects; if your work is more esoteric, it might not be nearly as good. I haven't tested much outside of my narrow domain.

ApatheticCosmos 7 hours ago||
I started using Claude right before 4.5 came out, and 4.6 is where it turned a corner for my use.

Qwen 3.8 Flash Next is there. 3.8 27b is fairly close.

I'm excited to see what Qwen 4 will bring.

I'm running on a 128GB Strix Halo for Flash Next and an Intel Arc Pro B70 (32GB) for 27b.

MattyRad 5 hours ago||
I just used 5.5 xhigh reasoning to make a massive implementation spec (for a vibey throwaway project/exploration, not anything important, burned 80% of the 5h window), now my Strix is in the process of implementing it.

I think there's probably low-hanging fruit to outsource reasoning from rote read/writes.... Just speculation though, I'm not a token optimization expert.

1-6 2 hours ago||
I hope we're coming to a plateau with the HBM/GDDR7/on-chip RAM hype and get back to normalcy with system RAM alternatives for the rest of us.
pilooch 7 hours ago||
My goto private setup, runs ~50t/sex on a dgx spark with sglang, nvfp4. Excellent model.
apitman 7 hours ago|
This is a meaningless metric without knowing how long the sex takes
swiftcoder 7 hours ago||
30 seconds at most
fsiefken 7 hours ago|
I wonder if a higher Qwen3.8-27b quant could beat or match these lower < 16/24/48/64G Qwen3.8-Flash Next quants given similar quality.

What speed are you willing the sacrifice to debug/program for more complex jobs faster?

Then there are also these quants; https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...

happycube 2 hours ago||
Maybe, but those higher quants would need a large GPU accessible memory space - and the obvious candidates such as DGX Spark and Strix Halo don't have the bandwidth to run 27B at high quality quickly.

With Flash Next you only have ~6B active parameters so you can toss experts up into VRAM and/or run them on a CPU if you have enough RAM and bandwidth.

latentsea 6 hours ago|||
The benchmark indicates the IQ3_XXS quant beats 27B. I've switched to that now and am ditching 27B. Genuinely better results so far.
zkmon 7 hours ago|||
I'm not going to knock off my 27B-Q_6 for this. Good to to experiment though.
xreborn 7 hours ago||
from my experience dense models like 27b suffer less from quantization compared to large MoEs
More comments...