Top
Best
New

Posted by tosh 19 hours ago

Qwen3.8-Flash-Next(qwen.ai)
https://imageat.com/models/qwen-3-8-27b-uncensored
662 points | 214 commentspage 4
NooneAtAll3 17 hours ago|
what's the deal with absent scrollbar on the website?
tristor 11 hours ago||
Looks like it errors out in LM Studio using the Unsloth quants, apparently the Unsloth team has already posted patches for llama.cpp to support this.
jedisct1 11 hours ago||
MLX quants for Apple M5 with 128 GiB RAM: https://huggingface.co/jedisct1/Qwen3.8-Flash-Next-oQ4e-128k
stefan_ 15 hours ago||
I think these "Flash" models are sort of an evolutionary dead end. Sure, there are some routine tasks and applications where they can be used. But for the actual novel development work? It's much better to run a big model at high power for 30 mins than watch the Flash model struggle for 2 hours and produce massive churn.

Same reason your phone has a few big CPU cores for real work, it's much better to "race to idle" than have an "efficient" core struggle. Shitty experience, shitty power efficiency.

jononor 13 hours ago||
If you have good feedback signals, like tests/benchmarks/etc, then it is potentially better to do multiple turns where model uses that to adjust code. Which might not need as smart a model.
wolttam 15 hours ago||
It depends how you use the models. These small models work great for developers who prefer to stay more in the loop, and only task the model with things that can really only be interpreted in one way.

Not to mention, they’re great for self-hosting and getting yourself to not be dependent on some API that can go down or be altered at any time.

Big models seem to mostly be good for pushing ahead the frontier - the smaller models tend to gain the frontier’s capabilities after only a handful of months anyway. Many are perfectly content remaining a few months behind the bleeding edge.

axegon_ 17 hours ago||
Aaaaaaaaaaaaaand dario meltdown on twitter in 3, 2, 1...
khangtong988 17 hours ago||
Yoh yoh I like It
postal6666 5 hours ago||
[dead]
myshapeprotocol 17 hours ago||
[dead]
skarz 18 hours ago|
do we really need breaking news about qwen posted every single day?
KronisLV 18 hours ago||
If there’s news, then yes. This is a pretty great new release for those still stuck on Qwen3.6 35B A3B if they have enough memory but don’t have super powerful compute.

I wonder if I could get this running through vLLM on 6x Nvidia L4 - the 3.6 worked great on 4 cards but sadly TP6 just isn’t a thing and I don’t have 8 cards available, maybe it’s gonna be okay with like TP2 and MTP. I have no idea at this time, probably need to test out what even might be possible.

NitpickLawyer 18 hours ago|||
This particular release is interesting because it's a preview of qwen4 architecture. And, while benchmarks are iffy, this is a direct comparison, by the same team, with qwen3.8-27b that was pretty well received for a local model.

This "next" release adds a new concept, first public release with n-grams, I think. And it's in a MoE size that is likely to be very fast and cheap to serve (faster than 27b for sure). It's also well suited for inference on alternative compute (i.e. sparks, macs, etc) so it's relevant to local users.

pseudony 18 hours ago|||
I and presumably quite a few others with AMD AI or Apple Mac platforms are very impacted by this.

:)

It is very relevant and for a certain group of us, far more impactful to our work the next month(s) than any blog post could be.

tosh 18 hours ago|||
this is a new architecture (foreshadowing qwen 4)

> trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board

https://x.com/Alibaba_Qwen/status/2092591393424515114

c16 18 hours ago|||
There are many topics, personalities and politicians we hear about daily who have no merit.

Qwen's advances do (currently) have merit.

dofm 18 hours ago|||
This actually is meaningful news, I think. Pretty wide audience appeal in the local LLM space too.
iAMkenough 18 hours ago||
yes there’s no shortage of online real estate