Top
Best
New

Posted by snehesht 9 hours ago

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s(github.com)
514 points | 255 comments
a11r 5 hours ago|
I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I'm running 4-bit quants on an RTX Pro 6000 rented for approximately $1/hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: https://www.reddit.com/r/BlackwellPerformance/s/FrKwk3GoDK
sfifs 41 seconds ago||
I've run DeepSeek V4 Flash on DS4 on single DGX and standard model weights on aDGX cluster. There was some degradation going to the hybrid 2 but quant but really not much.
Winfred-zz 35 minutes ago|||
I just ran a set of benchmarks, ninfer-3090-qwen3.8-27b (so mix of Q4 and Q5) vs strata-qwen3.8-flash-next-iq3_xxs (so Q3):

│------------------- │ Ninfer-3090 │ Strata

│ Code generation │ 52/78 (66.7%) │ 70/78 (89.7%)

│ Code completion │ 40/50 (80.0%) │ 44/50 (88.0%)

│ Total------------- │ 92/128 (71.9%) │ 114/128 (89.1%)

│ API failures------ │ 10 │ 5

- Ninfer generation: ~122 min total.

- Strata generation: ~142 min total.

So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t/s, this modified version can do 50 t/s (but it's extremely long in it's thinking, it just goes on and on.

This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).

Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t/s.

That strata has been optimized on my Oh My Pi conversations. So when I'm using it, it's probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.

nialv7 2 hours ago|||
ISTA-DASLab's IQ3_S quant is really good https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RC...
ricardobeat 2 hours ago||
What are you running it on? I'm getting mixed results using the IQ3_XXS quant which supposedly matches baseline, it feels significantly degraded.
robot_jesus 5 hours ago|||
Can you say more about where you're renting the RTX Pro 6000 for $1/hour?
a11r 2 hours ago|||
I am renting spot VMs from Nebius. I've tried a variety of other providers like Vast and Spheron. Vast worked well for renting 5090s but I like the large memory and pricing I'm getting at Nebius for RTX Pro 6000. The extra RAM really matters because I need PLE to offload the ngram to RAM.
_zoltan_ 4 hours ago|||
vast.ai?
NinjaTrance 3 hours ago|||
Just as curiosity, how long does it take to set the environment up and running?

Is it viable to start/stop it multiple times per day?

a11r 2 hours ago||
Yes, you can probably get the whole thing up and running in about an hour the first time. If you pause and restart, it takes about 15 minutes to load the models from disk into GPU memory, so budget for cold startup time.
KeplerBoy 2 hours ago||
How does it take fifteen minutes to read <100 GB into GPU memory? Shouldn't that be limited by SSD speed with everything slower than a minute being a terrible ssd?
boredatoms 57 minutes ago|||
It also depends on the runtime, vllm is unbelievably slow at model loading compared to llama.cpp
teaearlgraycold 2 hours ago|||
A lot of cloud platforms have terrible slow network storage. They also might need to compile the GPU kernels fresh as they might not have a persistent CUDA cache.
sail0rm00n 5 hours ago|||
$1/hr sounds great. Where are you getting it for those prices?
transcriptase 4 hours ago||
[flagged]
dang 3 hours ago|||
Could you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for.

If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.

transcriptase 2 hours ago||
Apologies. I had recently read (I think from someone on here) that they were in some discord where those sellers advertise. Whoever it was mentioned that the service was incredible value for what they got, but admittedly it was for toy projects, there was no written agreements, and uptime was something like 98%. I should have found the source and linked, and will try to do better overall in my replies!
xingped 4 hours ago|||
No? Per someone else's comment, presumably they're using vast.ai which does indeed list the 6000 for $1/hr
nickpsecurity 1 hour ago|||
That's cheap, too. Which hosting service are you using?
aatd86 4 hours ago||
$1/hour ? I need that deal as well.
trvz 4 hours ago||
Not the original commenter, but Upcloud has them at that price for spot instances: https://upcloud.com/pricing/
Jackson__ 3 hours ago||
I've just tested Strata on a simple 50 image vision benchmark. The task is to output the exact coordinates of a requested object. The result via Strata had a median error distance of 154.8 pixels, avg of 168.8. Running the exact same GGUF and vision adapter weights on llama.cpp gives me a median error of 46.5, avg 81.4.

To put that into perspective, here are some more numbers from other models via llama.cpp:

Median/Average

Qwen 3.5 9B BF16: 46.5 / 193.3

Qwen 3.6 35B Q4 K XL: 38.4 / 76.4

Qwen 3.5 122B Q3 K M: 32.9 / 68.6

The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.

I have done no further testing, as these results line up perfectly with my expectations.

Xenograph 9 minutes ago||
Is this test available somewhere? Would like to test it out on my models.
throwaway219450 22 minutes ago|||
What does SAM3 get on the same test set?
NamlchakKhandro 1 hour ago||
Tldr, strata is a waste of time.
AntiRush 1 hour ago||
I've been working on support for this model in ds4 on the RTX 6000 pro - it's been really great for my use cases. The ds4 q4 quant performs a lot better than other similar sizes that I've seen.

Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:

  Code: prefill 1,251 tok/s decode 255.26 tok/s
  Prose: prefill 1,251 tok/s decode 198.78 tok/s 
Most important for me, I can run 4 concurrent streams at 400+ tok/s.

https://github.com/fairfieldt/ds4

jacquesm 1 hour ago|
DS4 is an odd model. I have it working on way too many GPUs and yet for many tasks Qwen 3.8 will do much better. It also tends to loop, which is super annoying.

GLM5.3 runs on similar hardware and is much better so if you're going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?

Other than that, when you're done with that card...

snehesht 9 hours ago||
I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

https://huggingface.co/Qwen/Qwen3.8-Flash-Next

roscas 8 hours ago||
Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.

I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.

StumpChunkman 7 hours ago||
How much VRAM on your 3080? I've got an early 10gb model. I've been thinking of exploring local coding models, but everyone seems to use much better GPUs than I have access to. Yours is one of the first I've seen with maybe similar hardware on some level.
roscas 7 hours ago||
Yes, 3080 with 10GB, forgot to mention that.

Mine is at the moment writting some cpp code for some SBOM tests.

I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.

Oh I will run some other tests with hermes now because hermes is amazing too.

thatsabadlook 8 hours ago|||
Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me
hdjrudni 5 hours ago||
Not sure you understand the term 'surprisingly well'. It means 'better than expected'. I suspect they parent poster didn't actually expect to get >= 100 T/s.
jacquesm 1 hour ago|||
Speed is one thing, accuracy another. Have you benchmarked it against a reference? If so, what were the results? I tend to go for accuracy over speed because usually that means fewer round trips and fewer tokens wasted.
proc0 9 hours ago|||
Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
incognito124 8 hours ago|||
Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore
mickeyp 8 hours ago|||
I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

snehesht 8 hours ago||
This is interesting, thanks. - https://github.com/Neroued/ninfer
roscas 4 hours ago||||
I prefer https://ornith.ai/ornith_1_5.html to Qwen 3.8 not only because it is much faster on my hardware but better responses.

But this Qwen 3.8 Flash next coder is amazing running with Strata.

JokerDan 7 hours ago||||
Is this true for 27b Q4_K_XL vs flash next IQ3_S? I thought under Q4 models start quickly degrading?
gruturo 5 hours ago|||
While this is generally true, it's _a little_ less true the larger the model is.

Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.

Tade0 2 hours ago|||
To add to the other comment, there's also Ridge quantisation - the majority of weights are indeed Q3_x, but the most sensitive layers are FP8.
snehesht 8 hours ago|||
Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.
nicce 8 hours ago||
I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.
DoctorOetker 4 hours ago||
do LLMs tend to be homesick when not used in the same harness they sat in during some training phase?
Bnjoroge 1 hour ago||
iirc there was a sectionin Qwen’s paper where they talked anout how they post-trained flash or 3.8 to work just as well regardless of the harness or eval used. I think that used to be true but not sure if it is any longer
a11r 5 hours ago||||
We recently moved from 27B to Flash Next. The quality is superior for coding. Our workload is primarily well-defined coding tasks that need to be attempted a few times before the model gets it just right. FlashNext is also better at finding issues in generated code than Gemini 3.8 Flash.
swozey 1 hour ago||
I'm on m1 max 64gb and went from qwen3.8-27B back to qwen3.6-a35b. Is flash next the move? I went from usable say 40tk/s qwen3.6 to unusable, like 11 with 3.8 and not impressed with the replies for the time sacrifice. pi (omp) and omlx but not with the recent 3.8 patch.

I've been waiting for a 35b of 3.8, I don't really know what the other versions are about. I'm on 5g so juggling 40gb of model files sucks. And honestly I'm sick of tweaking this stuff for no, very little, or break-it level improvements. Qwen3.6-a35b has been solid for work, just don't give it freedom to wipe your data.

thatsabadlook 8 hours ago|||
Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.
geye1234 8 hours ago||
I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
thatsabadlook 1 hour ago|||
Definitely not.
PcChip 7 hours ago|||
Spelling mistakes?

What inference engine are you using for flash next?

geye1234 27 minutes ago|||
I'm running Pennyroyal's Docker image (on Podman) which uses sglang. I have a single RTX 6000 Blackwell and 128GB RAM. I turned off disk caching. I'm running with a ~500K context, but have been limiting it to 256K in the client (pi).

It always detects its spelling mistakes, btw, but it worried me. It may turn 'rm -rf ' into 'rm -rf /' one day.

Almost certainly the problem is my config, not the image.

anon373839 7 hours ago|||
Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.

Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)

notnullorvoid 7 hours ago||
Which quantization are you using to reach those numbers?
jacquesm 1 hour ago||
LLM threads the world over are spammed with Strata links, it remains to be seen how much of the breathless hype remains standing once the honeymoon period is over. I've tried it but so far I have not seen anything that overly impressed me in terms of accuracy, though the speed is definitely there. I'm sure there are applications for LLMs where the quality of the answers is less important but I don't have any of those. YMMV.
kamranjon 8 hours ago||
Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well.

https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...

lxe 6 hours ago||
Why isn't this type of expert caching in the native llama.cpp yet? Why do we need a separate codebase?
parsimo2010 6 hours ago||
One reason is that the llama.cpp team (GGML) has strict requirements that a human must understand the code they are contributing. If a project is fully vibe coded they can’t contribute. So a lot of projects where an AI went and coded a bunch of custom kernels to increase speed are left to their own devices.

I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.

sillyfluke 5 hours ago|||
>If a project is fully vibe coded they can’t contribute

The irony (however mild) is apparently lost on the rest of the field.

Loquebantur 5 hours ago|||
"A human pretends to understand it" signifies what exactly?

What you really mean is, the core team there doesn't want to lose control.

Which isn't really predicated on contributions not being "vibe coded" or whatever.

When quality is the problem, you need to be able to make your standards explicit, or you're just gatekeeping irrationally.

anamexis 5 hours ago|||
They do make their standards explicit: https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTI...

What part do you think is irrational gatekeeping?

rfgplk 4 hours ago||
> A proper code review usually takes something like one hour per 200-400 LOC and you should be spending at least that much time on code review alone.

Not only is this not enforceable (how do you enforce how long someone spent working on a codebase on their own local machine?) the metric is severely off which instantly makes me question the competence of the llama.cpp dev team. You can easily review 10-100x that in an hour, even if you're being super pedantic about it.

I also just ran _one_ of their files (with include deps) through Astra and it detected >100 vulnerabilities/correctness errors (with over 10 outright UB/memory corruption issues). It's actually outright shocking.

rpdillon 3 hours ago|||
> question the competence of the llama.cpp dev team

Following them for years, they seem extremely well put-together, and have excellent judgment. They are using the same policy as Linux and Debian (in my words, the speed of light is human understanding and judgment). Whether it is reasonable is a different question from enforcement, which typically comes down to "this seems fishy, explain your reasoning".

As for code review, the rule of thumb I've used for decades is: it takes about as long to review and understand as it does to write. Your 100x metric is completely outside of anything I've seen in any hobby or professional project, ever.

I'd like to see specific files you scanned and specific vulnerabilities cited.

calebkaiser 2 hours ago||||
If you can thoroughly and accurately review 100 * 400 = 40,000 lines of CUDA kernel code per hour, I know roughly 1,000 people who would love to hire you right now.
kube-system 3 hours ago||||
What you have quoted is a sentence elaborating on the requirements listed in the document. This is provided to help you better understand why the requirements exist and the goal they are trying to accomplish.

> should

https://www.rfc-editor.org/info/rfc2119/

The reason you SHOULD take that time to read the output is because you must read it to understand it.

And the way this is enforced is explicitly called out in the document (and again in more detail in the linked AGENTS.md): the maintainers may ask you to explain it.

jacquesm 1 hour ago||||
> You can easily review 10-100x that in an hour,

Frankly, I don't believe you. I'm half decent at writing CUDA directly (a holdover from a project a few years ago and it is a nice skill to have), the degree to which these are optimized is unlike 99.9% of all other code out there and even a tiny slip-up is either going to kill your results, your performance or both and if you're lucky only in some edge case. Understanding this code is hard work. I made a couple of minor edits to some .cu files in llama.cpp yesterday because I have a pretty weird setup which they obviously did not anticipate and it took a couple of hours to get it 'just so'.

IsTom 3 hours ago||||
> You can easily review 10-100x that in an hour

40k LoC per hour of pedantic review? That's eleven lines per second, every second, for an hour.

Luker88 3 hours ago|||
> You can easily review 10-100x that in an hour

200-400 LOC, 10-100x = 2.000-40.000 LOC/hour for human review?

reviewer: LGTM

Just merge in main, what are you even pretending to review?

AI review should happen before human review, not instead of it.

I see frontier AI giving up and finding only nitpicking things on huge PRs, then finding logic bugs that were always there after cleanup.

Split your PR in smaller ones, both humans and AI will work better.

rfgplk 4 hours ago|||
> What you really mean is, the core team there doesn't want to lose control.

It's 100% this. They basically produce vague guidelines such that only the core maintainers are allowed to use LLMs, under the guise of "well of course we understand the code" and no one else is. It's also completely unenforceable, how are they going to prove whether someone understands the code or not? Even if they show sufficient evidence/understanding the maintainers can simply sabotage them and accuse them of using an LLM to explain the code. No one wins here.

rpdillon 3 hours ago||
> how are they going to prove whether someone understands the code or not?

By discussing the code.

> maintainers can simply sabotage them and accuse them of using an LLM to explain the code

Bad faith enforcement is possible no matter the rules. If you think it's bad faith, a different policy won't save you.

hgoel 6 hours ago||
The "mainstream" inference engines are notoriously slow to integrate this stuff, to an extent understandably given the complexity of ensuring numerical accuracy alongside supporting a wide array of systems and models. Part of it is that not everyone is willing to bring what they develop into a pull request because they vibe coded it and don't care to deal with whatever quality requirements the more well known inference engines have.
mmaunder 7 hours ago||
More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.
Culonavirus 4 hours ago||
The real problem is the DRAM mafia and artificial scarcity. One of my notebooks is almost 3 years old, effectively similar spec now - same price (a bit higher actually). My desktop PC built around march/april 2023 (4090, 64gb ram, 7900x3d) is now pretty much still the top dog out there due to the gpu and fast ram insanity and if I wanted to sell it today, I'd get more money for it now used and over 3 years old that when I bought it!

We should be having 64/72+ GB video cards by now. 128GB+ system ram prosumer laptops and 256GB+ system ram prosumer/gamer desktops. But it all went to shit and it will require some brutal datacenter and datacenter-adjacent bankruptcies before it gets better.

Some of these greedy bastards need to lose their pants on all of this.

latentsea 5 hours ago||
You can run IQ3_XXS, IQ3_S and the IQ4_XS quants on this too. It works. It's fantastic. I'm getting better results than 27B now.
bitexploder 5 hours ago||
To add: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RC... IQ3_XXS is within a point of the fully unquantized model and IQ3_S actually beats the unquantized model on many tasks! You lose absolutely nothing. It is quantization magic :)
SuperV1234 7 hours ago||
We're getting closer and closer to the day we can have an Opus-like model running locally. The dream!
stymaar 6 hours ago||
It depends on what you mean by “Opus-like”, because if you mean “as strong as Opus 4.6 for agentic coding” then Qwen3.8-27B has been there for the past two months.

But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).

nycdatasci 6 hours ago|||
Opus 5.5 was a step function change. Really crushes on multi-hour coding compared with prior models.
jjcm 4 hours ago||
Highly agree.

I keep saying “I’d be so Happy with ${currentOpusVersion} locally”, but I keep being impressed with how much the capabilities change between versions. I have a RTX 6000 pro so I can easily run this qwen 3.8 flash next, but it’s much harder to give up the freedom that 5.5 gives me.

hgoel 6 hours ago||||
The most visible leap between 4.6 and 5.5 seems to be that the latter has gotten much more computer-use training, so there's a clear progression in the ability for the model to use Blender. But catching up on that is just a matter of training on the same thing.
system2 6 hours ago|||
I would be forever happy with Opus 4.8.
bitexploder 5 hours ago||
Flash next is /really/ close. It is at parity with 4.7 as far as I can tell and basically where Opus 4.8 was. It is a genuinely good model. And I run it at home on $1500 of GPU at 125 t/s :)
gruturo 5 hours ago|||
Seconding this. Flash-next (and let's not forget, it's a PREVIEW of the 4 architecture - with the "real" 4 rumored coming later this month) is the first model I can run on reasonable hardware (2 thoroughly obsolete P100s off ebay at ~$100 each plus the RAM I could scavenge from other PCs at home) at a reasonable speed (22, with GPUs in layer-split due to llama-cpp's limitation on qwen4-exp arch, and no MTP. Strata could double these numbers).

It's.... the real thing, for the first time. If you cut me off cloud models today, I would get plenty of utility out of this thing.

(Others may have had the same feeling from GLM5.3 or Deepseek 4.1 flash but I never had a chance of running those.)

fsiefken 2 hours ago|||
Running an nvidia card at full load, would cost me ~100 euro of electricty each month (europe). Of course one wouldn't have usage caps.

Where the internet was a subscription 15 euro subscription to encyclopaedic knowledge, an genAI subscription is renting a researcher/programmer for 100 euro.

stymaar 2 hours ago||
> (europe)

Given the massive difference in electricity price between different european countries, adding "Europe" doesn't bring much context.

magicalhippo 1 hour ago||
Even with a country. Here in Norway we have multiple price zones, and at times there can be 100x difference between them, often 10x. All due to lack of transmission capacity between northern and southern zones.
latentsea 5 hours ago|||
We are already there. Prior to Strata the best I could run was Qwen3.8-27B at Q6, which itself is already at like Opus 4.5/4.6 level, and now with Strata on an R9700 and 64GB of RAM I can run Qwen3.8-Flash-Next IQ3_XXS at 60 t/s. It's even better. You can run it on even more modest hardware with Strata too.

Plus they announced Qwen4-Flash. It's not released yet, but it's the same architecture as Qwen3.8-Flash-Next, which now runs fast on consumer hardware.

Opus at home is a thing now.

konaraddi 4 hours ago||
Another 2-5 years from now, it may even become accessible to most people (current barriers being cost and technical expertise, bottleneck is cost).
trvz 3 hours ago||
The mentioned “Qwen3.8-27B at Q6“ can be run on a <1500$ Mac mini and with the barest technical expertise.
bitexploder 5 hours ago|||
That is right now. This model is easily as good as Sonnet5 / Opus 4.7 on DeepSWE. I have benched over half of DeepSWE now on a 3 bit Flash Next quant and it is at parity with Sonnet and Opus 4.6/4.7. It finishes most of the tasks they do. Overall it is within 1 point.

FWIW Qwen 3.8 27B is just slightly behind and basically Sonnet 5 high. I have been benching these models. We have Opus at home. :)

copx 6 hours ago||
Dream or nightmare?

In face of the recent Hugging Face incident we should really be concerned about the security implications.

What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.

We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..

nvme0n1p1 6 hours ago|||
https://huggingface.co/blog/security-incident-july-2026

- OpenAI hacked Hugging Face

- OpenAI models refused to help Hugging Face during incident response

- Hugging Face turned to GLM, who helped in the defense

That pattern repeats over and over. https://www.felonybench.com/

You should be happy that open weight models exist. They're the last thing protecting the internet from the unconvicted felons working at OpenAI+Anthropic.

r14c 6 hours ago||||
Secure systems are possible, but now we have a compelling reason to actually write them. Everything can be trivially hacked because the industry is pathologically adverse to security being part of the design process.
lxe 6 hours ago||||
Nothing is stopping it. This model is woefully bad at accurate creative red teaming however. GLM finetunes on the other hand are pretty good. And I'd bet they are already deployed and doing all sorts of deeds.
coursenumpls 6 hours ago||||
if the alternative is all human intelligence is cucked by 2-3 amoral American labs then we've had a good run, don't care.

my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.

4858585858 6 hours ago|||
[flagged]
Tepix 8 hours ago|
Q2 quantization. Not interested.
tcdent 7 hours ago||
All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.
bitexploder 5 hours ago|||
I am benching Flash next on a 3 bit XXS quant and it is holding just fine against published benchmarks. Using DeepSWE official harness and Pi with absolutely zero benchmaxx or harness config. Install stock Pi and running my agents in it. I am halfway through DeepSWE (it takes FOREVER, even at 125 t/s) and it is neck and neck with Opus 4.7 and Sonnet 5.

On a 3 bit quant btw.

I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)

It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)

apitman 2 hours ago||
How much VRAM total you using for this? I have a bunch of 3060s in a threadripper and thinking I might need to give Flash Next a try. Currently using 27B.
bitexploder 40 seconds ago||
64GB across 2 cards. You need the MoE caching build and enough regular RAM to fit everything else to get decent performance. You should get reasonable performance if you have enough RAM. I have a 4080 that does around 40 t/s right now on a MoE caching build. It's great. Not fast, but chugs along.
latentsea 5 hours ago|||
You can run a 4 bit quant with this. Personally, I switched to running IQ3_XXS and am getting better outputs than 27B and at faster speeds.
ivanjermakov 6 hours ago|||
These "revelations" are getting closer and closer to "download RAM for free" each day.
latentsea 5 hours ago|||
You can run IQ3_XXS, IQ3_S, and IQ4_XS too. I've switched to IQ3_XXS and am running at 60 t/s on Strata vs the 21 t/s I was getting in llama.cpp. Better outputs too.
CamperBob2 5 hours ago||
Larger models can tolerate Q2 quantization surprisingly well, especially if they were trained with quantization in mind. I don't know about 3.8 125B, but for example, there are 2-bit quants of Kimi K3 that exhibit strong reasoning and maintain decent coherence at longer contexts.
More comments...