Top
Best
New

Posted by erdaltoprak 9 hours ago

Qwen 3.8 27B(huggingface.co)
833 points | 548 commentspage 7
walrus01 2 hours ago|
Is it just me or is 3.6 27B Q8 K XL (Unsloth) holding up better in sustained token/s rate as the context fill increases over time? The token/s rate seems to be much higher for a time period deeper into context than previously seen.

At least as compared to 3.6 27B in the same quantization.

arjie 8 hours ago||
I use the Qwens as a vision model for my DeepSeek V4 Flashes to handle. But the Qwens run on old RTX A6000 Ampere. Does anyone know if there's any news about INT4/AWQ quants for the RTX A6000?
ericd 8 hours ago|
Was recently thinking about doing something similar, do you basically just have the qwens describe what they see for the flashes?

Was considering adding a LoRa/vision head to Flash, but seems like it could take a while to get it right.

If DSv4 Flash was multimodal, I’d probably be done model shopping for a while

arjie 8 hours ago||
Same, with a multimodal DSv4 Flash I would just stop paying attention to things. Very smart, and at 260 tok/s it's too fast to care about anything else. If you ever graft something like that I would love to hear about it.

Yes, I have a very dumb flow. The harness has a describe_image tool that takes an image and a prompt and so DSv4 Flash uses it to get an idea of what it's looking at.

ericd 7 hours ago||
Yeah, I might just replicate what you're doing. Main issue right now is just finding spare vram to actually run another model in parallel... And yeah, if I train up a vision adapter somehow, I'll try to put it up/post about it, seems like we're getting the killer apps for local LLMs right now, where it's just feasible enough if you're enthusiastic enough to be a bit economically irrational, and just useful enough to sort of rationalize.
yassa9 9 hours ago||
Can anyone who has that specific personal test he tries on different models , and tries this model , to tell us here if possible , how good or bad is this new model ? compared to others ?

I only trust those users genuine personal tests

alyandon 9 hours ago|
There is a down to earth guy on YT that performs a series of tests against LLMs running on non-god-tier commodity hardware. He will likely be testing this soon enough.

https://www.youtube.com/@lukesdevlab

I don't know if that is what you are looking for or not and as always your experiences may be different.

xscott 6 hours ago|||
So much potential for that channel. He's got a nice range of tests and a no nonsense presentation style.

However, watching tests of heavily quantized models that weren't designed for it (non-QAT) is frustrating. There's no way to tell if the actual model fails because it's dumb or if the lobotomy made it that way.

alyandon 5 hours ago||
I noticed he does pay attention to feedback on his videos and I think some people have pointed that out.
yassa9 8 hours ago|||
thaaanks man, this channel seems really informative, although < 10K subs only !
alyandon 8 hours ago||
It's a relatively new channel - but yeah - I feel the guy puts a lot of effort into what he does and deserves more subs.
esotericsean 6 hours ago||
Need to upgrade to a second 3090! Slowly building up my local models with Krea2, MiniMax H3 (and their new Music3), and now Qwen 3.8
ThouYS 9 hours ago||
3.6-27B on little-coder was already mind blowing. looking forward to this guy!
jlkivey 9 hours ago||
Note: on the model card the comparison to Opus is Opus 4.6 Max, not 4.7
irthomasthomas 8 hours ago||
Why don't qwen/alibaba host the model themselves? I was looking forward to trying it on their coding plan. Google are the same way with their Gemma models.
spwa4 8 hours ago|
Pretty sure you can use Gemma models on Google's "Vertex AI".
crazyemeraldcod 2 hours ago||
Its so smart!
tosh 9 hours ago||
also cool: Qwen 3.8 27b is multi modal!
gurkwart 9 hours ago|
strong visual reasoning apparently, which is nice. still lacking native audio however. hoping for more companies to embrace the spirit of something like `gemma-4-12b-qat` for actual multi-modality (text, image, video, audio).
anana_ 9 hours ago|
Monstrous benchmarks! Hoping it is not benchmaxxed.
sheepscreek 4 hours ago|
I thought the same. But why claim something so shocking when it can easily be discredited and puts your reputation at risk? If they’re claiming Opus 4.6 level, I expect it to at least match Sonnet 4.6.
More comments...