Posted by snehesht 9 hours ago
│------------------- │ Ninfer-3090 │ Strata
│ Code generation │ 52/78 (66.7%) │ 70/78 (89.7%)
│ Code completion │ 40/50 (80.0%) │ 44/50 (88.0%)
│ Total------------- │ 92/128 (71.9%) │ 114/128 (89.1%)
│ API failures------ │ 10 │ 5
- Ninfer generation: ~122 min total.
- Strata generation: ~142 min total.
So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t/s, this modified version can do 50 t/s (but it's extremely long in it's thinking, it just goes on and on.
This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).
Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t/s.
That strata has been optimized on my Oh My Pi conversations. So when I'm using it, it's probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.
Is it viable to start/stop it multiple times per day?
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
To put that into perspective, here are some more numbers from other models via llama.cpp:
Median/Average
Qwen 3.5 9B BF16: 46.5 / 193.3
Qwen 3.6 35B Q4 K XL: 38.4 / 76.4
Qwen 3.5 122B Q3 K M: 32.9 / 68.6
The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.
I have done no further testing, as these results line up perfectly with my expectations.
Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:
Code: prefill 1,251 tok/s decode 255.26 tok/s
Prose: prefill 1,251 tok/s decode 198.78 tok/s
Most important for me, I can run 4 concurrent streams at 400+ tok/s.GLM5.3 runs on similar hardware and is much better so if you're going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?
Other than that, when you're done with that card...
This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.
I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.
I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.
This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.
Mine is at the moment writting some cpp code for some SBOM tests.
I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.
Oh I will run some other tests with hermes now because hermes is amazing too.
It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.
But this Qwen 3.8 Flash next coder is amazing running with Strata.
Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.
I've been waiting for a 35b of 3.8, I don't really know what the other versions are about. I'm on 5g so juggling 40gb of model files sucks. And honestly I'm sick of tweaking this stuff for no, very little, or break-it level improvements. Qwen3.6-a35b has been solid for work, just don't give it freedom to wipe your data.
What inference engine are you using for flash next?
It always detects its spelling mistakes, btw, but it worried me. It may turn 'rm -rf ' into 'rm -rf /' one day.
Almost certainly the problem is my config, not the image.
Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)
https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...
I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.
The irony (however mild) is apparently lost on the rest of the field.
What you really mean is, the core team there doesn't want to lose control.
Which isn't really predicated on contributions not being "vibe coded" or whatever.
When quality is the problem, you need to be able to make your standards explicit, or you're just gatekeeping irrationally.
What part do you think is irrational gatekeeping?
Not only is this not enforceable (how do you enforce how long someone spent working on a codebase on their own local machine?) the metric is severely off which instantly makes me question the competence of the llama.cpp dev team. You can easily review 10-100x that in an hour, even if you're being super pedantic about it.
I also just ran _one_ of their files (with include deps) through Astra and it detected >100 vulnerabilities/correctness errors (with over 10 outright UB/memory corruption issues). It's actually outright shocking.
Following them for years, they seem extremely well put-together, and have excellent judgment. They are using the same policy as Linux and Debian (in my words, the speed of light is human understanding and judgment). Whether it is reasonable is a different question from enforcement, which typically comes down to "this seems fishy, explain your reasoning".
As for code review, the rule of thumb I've used for decades is: it takes about as long to review and understand as it does to write. Your 100x metric is completely outside of anything I've seen in any hobby or professional project, ever.
I'd like to see specific files you scanned and specific vulnerabilities cited.
> should
https://www.rfc-editor.org/info/rfc2119/
The reason you SHOULD take that time to read the output is because you must read it to understand it.
And the way this is enforced is explicitly called out in the document (and again in more detail in the linked AGENTS.md): the maintainers may ask you to explain it.
Frankly, I don't believe you. I'm half decent at writing CUDA directly (a holdover from a project a few years ago and it is a nice skill to have), the degree to which these are optimized is unlike 99.9% of all other code out there and even a tiny slip-up is either going to kill your results, your performance or both and if you're lucky only in some edge case. Understanding this code is hard work. I made a couple of minor edits to some .cu files in llama.cpp yesterday because I have a pretty weird setup which they obviously did not anticipate and it took a couple of hours to get it 'just so'.
40k LoC per hour of pedantic review? That's eleven lines per second, every second, for an hour.
200-400 LOC, 10-100x = 2.000-40.000 LOC/hour for human review?
reviewer: LGTM
Just merge in main, what are you even pretending to review?
AI review should happen before human review, not instead of it.
I see frontier AI giving up and finding only nitpicking things on huge PRs, then finding logic bugs that were always there after cleanup.
Split your PR in smaller ones, both humans and AI will work better.
It's 100% this. They basically produce vague guidelines such that only the core maintainers are allowed to use LLMs, under the guise of "well of course we understand the code" and no one else is. It's also completely unenforceable, how are they going to prove whether someone understands the code or not? Even if they show sufficient evidence/understanding the maintainers can simply sabotage them and accuse them of using an LLM to explain the code. No one wins here.
By discussing the code.
> maintainers can simply sabotage them and accuse them of using an LLM to explain the code
Bad faith enforcement is possible no matter the rules. If you think it's bad faith, a different policy won't save you.
We should be having 64/72+ GB video cards by now. 128GB+ system ram prosumer laptops and 256GB+ system ram prosumer/gamer desktops. But it all went to shit and it will require some brutal datacenter and datacenter-adjacent bankruptcies before it gets better.
Some of these greedy bastards need to lose their pants on all of this.
But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).
I keep saying “I’d be so Happy with ${currentOpusVersion} locally”, but I keep being impressed with how much the capabilities change between versions. I have a RTX 6000 pro so I can easily run this qwen 3.8 flash next, but it’s much harder to give up the freedom that 5.5 gives me.
It's.... the real thing, for the first time. If you cut me off cloud models today, I would get plenty of utility out of this thing.
(Others may have had the same feeling from GLM5.3 or Deepseek 4.1 flash but I never had a chance of running those.)
Where the internet was a subscription 15 euro subscription to encyclopaedic knowledge, an genAI subscription is renting a researcher/programmer for 100 euro.
Given the massive difference in electricity price between different european countries, adding "Europe" doesn't bring much context.
Plus they announced Qwen4-Flash. It's not released yet, but it's the same architecture as Qwen3.8-Flash-Next, which now runs fast on consumer hardware.
Opus at home is a thing now.
FWIW Qwen 3.8 27B is just slightly behind and basically Sonnet 5 high. I have been benching these models. We have Opus at home. :)
In face of the recent Hugging Face incident we should really be concerned about the security implications.
What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.
We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..
- OpenAI hacked Hugging Face
- OpenAI models refused to help Hugging Face during incident response
- Hugging Face turned to GLM, who helped in the defense
That pattern repeats over and over. https://www.felonybench.com/
You should be happy that open weight models exist. They're the last thing protecting the internet from the unconvicted felons working at OpenAI+Anthropic.
my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.
On a 3 bit quant btw.
I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)
It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)