Top
Best
New

Posted by snehesht 12 hours ago

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s(github.com)
571 points | 272 commentspage 5
lousken 8 hours ago|
lm studio bionic, unsloth, now this... it would be nice if it worked at least in one of those without installing another component
kennywinker 6 hours ago|
I'm sure one of those inference engines will add support for this within a month or so.
quietFalcon 11 hours ago||
Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
snehesht 11 hours ago||
They have some community benchmarks published https://github.com/Niko1221/Strata/tree/main/bench/results
merbanan 11 hours ago||
Q2_0 does 33 tok/s decode and ~600t/s prompt processing at 128k context on RTX2060 8GB VRAM.

ISTA IQ3_XXS does ~21 tok/s decode and ~240t/s prompt processing

cesarvarela 7 hours ago||
It's funny that most of the AI industry is built around the assumption (which is most likely true) that it is not possible to run SOTA models on current consumer hardware.

Imagine if someone managed to run an Astra- or Fable-level model on a 5090 at reasonable speeds.

kennywinker 6 hours ago|
It's not crazy to imagine something like that, but something's gotta change before it can happen.

Either the 5090 part - new hardware that's tuned for AI specifically. But we won't see that until the datacenter buildout collapses or finishes, since they are buying up all of TSMCs capacity.

Or perhaps it comes from the model. 1-2 years ago it would be inconceivable to use a 27b model for coding and expect any kind of usable results. Today, I have a model that feels like it crosses the threshold from a toy to a tool, and i can run it on dated pro-sumer hardware. I don't think we'll ever see SOTA on consumer hardware, but as the small models cross more and more thresholds the gap will matter less and less.

0xbadcafebee 11 hours ago||
Lol, sure, if you quant it to hell (Q2) it'll go real fast...

They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.

It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.

snehesht 10 hours ago||
You're right but for simple use cases its useful. Someone pointed to ds4 + qwen3.8 with Q4_K will try that out.
sigbottle 10 hours ago||
It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?
nottorp 9 hours ago|||
Is Qwen 3.8 at Q4 good enough?

I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.

cycomanic 7 hours ago|||
I've run 3.8 flash next k4_xl on my Strix halo box (128GB). And in the work I have done so far it was not significantly worse than recent GPT (running default model on pro plan). Admittedly I was not doing complex work (reorganizing a jupyterbook), but I could not see significant difference in the quality of the work. It was a striking difference to Laguna s 2.1 which I had tried just before (much faster and much better quality).
nottorp 6 hours ago||
At Q what? That's what I'm mostly asking about.

My little test was "generate me a single page tic tac toe game in plain javascript. computer always play O. add unbeatable minmax. have the board, a status line and a new game button'. I used both lm studio and whatever the name of their new coding assistant that supercedes lm studio is.

Qwen 3.5 Q4 went into some kind of loop where it fixed whatever was broken on the previous iteration only to have it broken some other way. (I was writing the description of the errors).

Paid $20/mo claude opus did it right the first time. Or at worst it fixed the code based on descriptions without entering a breakage loop, iForgot. I know it isn't fair because it has 1 million tokens but still, it was just tic tac toe.

But since everyone says qwen is decent, it's either:

- Q4 is too little

- my idea of "decent" is too much

- 3.8 is much better than 3.5 even at Q4

Zambyte 5 hours ago|||
3.8 is a huge step up from 3.5, quantized or not. I do all of my programming on Qwen 3.8 27B Q4 these days.
MaxikCZ 10 hours ago||||
New models are trained with 8/4bit quantization in mind. Going from "native" 8 to 4 isnt as big of a step as going from 8 to 4 if native is full bf16.
amelius 10 hours ago|||
3 is the magic number, and 4 > 3.

(seriously, nobody knows why any of this works; it's just a matter of trying)

boredatoms 3 hours ago||
It hard to take below-8bit quants seriously
panny 11 hours ago||
I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
somenameforme 10 hours ago||
The card in question here had an initial MSRP of $1600. It's been bumped up by the market, probably because it turned out it's nice for things like this, but it's hardly in the 99% can't afford domain, especially if you're using it to replace a never-ending rent at which point it will pay for itself very rapidly, especially for heavy LLM users.

In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.

the__alchemist 9 hours ago|||
It was available at this MSRP 3 years ago, direct from Nvidia. It now goes for 3-4k USD, as you point out. MSRP stopped being a useful value for graphics card around that time. You realize the price hike, but still mentioned that 1.6k figure after.

No normal person is spending 3-4k on a GPU from 3 years ago. The availability is also of questionable provenance.

somenameforme 8 hours ago||
The MSRP is a good proxy for the 'level' of a card. It's not like the 4090 was a freak outlier. Cards of a comparable price offer comparable performance. So the 'level' for running a frontier level model at high performance is now at $1600 and continuing to trend sharply downwards.

Another nuance is that the computer hardware market is currently extremely inefficient in a way I don't understand. You can pick these cards up locally at places throughout Asia for around $2k new. That's retail single unit prices. No idea what's stopping somebody from closing the gap and making a ton of money - perhaps tariffs and data centers purchasing in a price insensitive fashion. Whatever the exact reason may be, what people pay for hardware is increasingly just radically different depending on where you buy it at.

the__alchemist 7 hours ago||
That arbitrage opportunity is remarkable. Maybe duty/taxes, as you say?

I'm the guy who (Who plays games and runs molecular dynamics simulations and other CUDA stuff) said 3 years ago "$1600 for a graphics card? That is excessive. I'll upgrade in a few years when ready" And bought a 4080 for $1200 from Nvidia instead of the 4090. Oops! Now there is no reasonable upgrade path.

panny 7 hours ago|||
>The card in question here had an initial MSRP of $1600.

63% of Americans can't come up with $400 in an emergency.

https://www.investopedia.com/here-s-how-many-americans-can-t...

The richest country in the world. Where all 50 states consume more than any other country in the world.

https://x.com/cremieuxrecueil/status/2102889196000256219

Can't come up with %25 of that in an emergency. (Even though the real price is something like 2-3x more than MSRP)

It must be nice, up there where you are so incredibly disconnected from reality.

MaxikCZ 10 hours ago|||
The reason models running on low vram are not talked about enough is because they are just not worth it. Qwen3.8 27b changed that, but even 24gb vram is too low for it. Running better model faster at 12gb vram is where its now at, and thats why you see people talkin about it
liuliu 9 hours ago|||
Because that's not possible (to have a GPT 5.6 Sol level model). People won't believe this and will keep dreaming, but intelligence is not free and 8GiB (shared with OS and other processes) is too small to be useful. Whether it is possible for 48GiB or 64GiB (meaning useful for model would be ~16GiB to 24GiB) with external fast storage (SSD), OTOH, is a question mark.
MrDrMcCoy 11 hours ago|||
Ternary Bonsai 2 might be for you.
luke-stanley 10 hours ago||
I might try running the expert pruned Coder model but yes, that PrismML Bonsai 2 Ternary 27B model is from the Qwen 3.8 27B model, which has better intelligence density (Artificial Analysis says), without the MoE disk use or architecture complexity (if you care about that)! There are also DFlash 2 models for it too (though in my experience this only measured faster for parallel requests, but I have a 3090). I am curious about the phone acceleration for Bonsai 2!
throwawayffffas 9 hours ago||
The 1% can afford to run these models without quantization.
lsb 8 hours ago||
There’s other slop projects to run of Qwen, like ds4, would be interesting to see a comparison
happycube 4 hours ago|
V4.0 Flash(-vision) in released form is "only" ~180GB of FP4 weights, so a braindead quant could work in 128GB of RAM.

4.1 is much larger, even leaving out the PLE.

api 10 hours ago||
Continued progress on these fronts is another reason I think the data center buildout is a bubble. It posits that AI use and growth will require an ever-increasing amount of power and floor space, which contradicts the entire history of computing. The high cost of data centers is largely electricity and floor space, which means there's a huge forcing function to make both the silicon and the software more efficient.
MaxikCZ 8 hours ago|
> Large costs of datacenters is largery electricity and floor space

Really? I would guess that those would be almost a rounding error on the price of gpus sitting in there

api 5 hours ago||
That too, and software efficiency directly attacks that.
deadbunny 11 hours ago||
> Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

And I thought piping to bash was bad

Skunkleton 9 hours ago||
I've never understood the security argument people are making when they complain about `curl foo | bash`. I get that these scripts sometimes mess up your bashrc or whatever, but from a security perspective I see no issue. You are already installing software from the same domain. If they were going to do something nasty, they could do it with any of the software you are using from them. It doesn't have to be the setup script.
minitech 9 hours ago|||
To compare it to just one other option: when you run `npx foo`, you know* that you’re getting the same public artifact that anyone else running it at the same time would get. (If you have a `min-release-age` configured, you also benefit from that.) If I wanted to distribute software like this, I’d include npm-shrinkwrap.json; then, with `npx foo@1.2.3`, you could be similarly confident in getting the same app every time.

(I picked this option for ease of comparison, getting a couple of major security wins with very low effort; I don’t recommend `npx`ing stuff in an otherwise unprotected environment either.)

* well, you can be somewhat more sure

athrowaway3z 8 hours ago|||
So to be clear; the solution is then something like

`curl https://raw.githubusercontent.com/my/domain/setup.sh | sh`

Note we dont even have a hash there - just a promise that a third party (github) has a log of whatever was hosted at that url.

nagaiaida 7 hours ago||
maybe then we'll pin hashes instead of filenames, and oops now anybody can fork my/domain and hand out a link that looks official with whatever contents they wish
slowin 7 hours ago|||
While I don't think piping curl into bash is the most secure, npm installing has proven time and time again to open yourself up to supply chain attacks. At least with curl you know that you're getting the supply chain put together by the software author. With npm, every single library is a vector for attack every time you update.
minitech 7 hours ago||
You’d be getting the supply chain put together by the software author in either case. It’s common to do both badly, but if you care about doing it well, that’s easier with npm (e.g. shrinkwrap, as mentioned) and can be taken farther (the non-varying artifact thing).
slowin 7 hours ago||
Npm dependencies resolve at install time though, so some of the pinned versions may have changed ownership or otherwise been modified since the author pinned them. With the curl | bash solution you're getting a singular supply chain packaged by the author at release time.
cpuguy83 6 hours ago||
With curl|bash you are literally getting anything that happens to be in that script. These are frequently poorly constructed, so not check hashes or pin dependencies they install. Even security companies (see trivy supply chain attack) get these badly wrong.

I'm not replying here to say one is better than than the other (npm has obviously had its share of problems) but rather to combat claims that curl|bash is somehow safer, it absolutely is not, in fact it's all the bad stuff about npm without the pretense of being potentially safe.

ffsm8 9 hours ago||||
You can detect the use of curl|bash server side, hence it's an essentially undetectable attack vector. People have shown poc attacks of that kind all the way back in the 2010s

https://news.ycombinator.com/item?id=17636032

The original blog is no longer available though.

But I've not had that stop me from doing that myself, I am more towards the "I like easy" then the "I want to be secure" crowd

wsc981 8 hours ago|||
There’s a snapshot on web archive:

https://web.archive.org/web/20250109045029/https://www.idont...

throooooo 6 hours ago|||
How would you do that? Just rely on the user agent? I use curl quite a lot, but don't pipe to a shell.
sspiff 9 hours ago||||
The setup script often runs privileged (by calling sudo) and that's not unexpected when installing new software.

When I install something, and it asks for my root password later, I will be much more likely to think "hold up, this ain't right".

minitech 9 hours ago|||

  cat >> ~/.bashrc <<'EOF'
  sudo() {
    sudo install-drivers-without-your-permission
    command sudo "$@"
  }
  EOF
(this is not an endorsement of curl | sh, just an indictment of the state of software)
nagaiaida 7 hours ago||
when people left their laptops unlocked, we used to wrap their sudo so that the output would be pre- and postfixed with ascii dolphins
jeremyjh 9 hours ago|||
Can you give me a popular example that requires sudo? I don't think that is very common at all.
serf 9 hours ago||
every single bash replacement for one.

oh my zsh is a specific example.

chsh requires sudo on most installs.

spiorf 9 hours ago||||
People with less experience normalize that behaviour and when the domain is not trusted the habit let their guard down. See all the clickfix attacks.
layer8 9 hours ago||||
I push binaries from untrusted sources through VirusTotal before running them. Piping a Bash script from curl bypasses that. Furthermore, such Bash scripts, when they aren’t self-contained, make security checks more difficult than a self-contained archive, installer, or binary, even when downloading the script without immediate execution.
parsimo2010 8 hours ago|||
You could always curl the install script, and modify it to run the virus scan in between the build and install steps.
Iolaum 9 hours ago|||
Nothing is stopping anyone from pointing their agent to that script to review and audit it before running it.
layer8 9 hours ago||
I don’t believe an agent can do that effectively without a sandbox to run the script in, if the script isn’t self-contained.

And everyone running a research agent on every download can’t be the solution. It’s much more effective to crowdsource a security database based on hashes. But for that, the downloads need to be self-contained.

bee_rider 7 hours ago||||
The intended workflow is to download the install scripts, download the source code, read them both, and then start running things. That’s how Open Source is secured. Piping from bash to curl is just the most obvious warning flag.
majorchord 7 hours ago||
And practically zero people are actually using this "intended workflow" in the real world.
bee_rider 6 hours ago||
Yes, the status quo is quite bad, which is why it gets complained about a lot.
rlpb 8 hours ago||||
It's `curl foo | sudo bash` that's the bigger objection. Running software usually shouldn't require root, and then the equivalence argument you make doesn't hold.
serf 9 hours ago||||
a script isn't getting hashed to see whether or not it's the one the website intended to serve you, for one.

what use is hashing every piece of software that goes thru the distros package manager just to throw caution to the wind at the layer above it?

w.r.t. "it's already from the same domain" , well most bash/z install scripts either invoke a package manager or they download and untar a package that has nothing to do with the host domain, anyway.

thomastjeffery 9 hours ago||||
The real problem is that we just aren't using package managers. We should be using package managers. Package managers are really really good.
IshKebab 6 hours ago|||
There is no argument. It's just people's reflex reactions.

The technical excuses they come up with (e.g. that the server can detect it and send different content) are just post-hoc justifications for their instinct.

Just ignore them.

gchamonlive 11 hours ago|||
Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs.
snehesht 11 hours ago|||
Yeah, I was surprised at first then had to dig through setup.py and setup.sh files to figure out.
mrinterweb 7 hours ago||
And yet people will let AI agents run autonomously on their machines. I feel like we're reaching peak YOLO with security.
CurbStomper4 1 hour ago|
[dead]
More comments...