Top
Best
New

Posted by erdaltoprak 7 hours ago

Qwen 3.8 27B(huggingface.co)
702 points | 446 commentspage 5
syntaxing 5 hours ago|
Would I be surprised there’s bench maxing happening? Yes. But some users also use Q4 quantized and complain how dumb local models are.
g023 3 hours ago||
All this performance at such small model sizes, why are the API fees so high for the AI monopolists on this side of the world?
theanonymousone 7 hours ago||
I'm wondering whether any provider can offer this for cheaper $/token than the new DSv4 Flash, which is both cheaper and smarter :/

Completely local use is a different story, of course.

mraza007 6 hours ago||
Man what a week, We just had GLM 5.3 that came out and then we had smaller local model Qwen3.8-27B from Qwen

Just tried using Pi Agent and looks very promising

maherbeg 4 hours ago||
Does anyone have a https://tenstorrent.com/hardware/cards to try it on?
TomGarden 7 hours ago||
Really excited to see what people do with this. 3.7 27B was probably the best compromise between size and intelligence to run on consumer hardware
geek_at 2 hours ago|
Do you mean 3.6 27b? Because qwen 3.7 didn't have an open weight version
synergy20 7 hours ago||
I wish this can run directly on my RTX 4090, seems like 30B is the sweet spot for dense model to run locally, sadly RTX 5090 is very expensive and I need a new PC and new power supply(and UPS) to run that, adding a second RTX 4090 is another option, but not sure if my PC can do that yet.
baron3dl 7 hours ago||
even a 3090 will give you the VRAM headroom. i run Q8 on an 3090/A6500 combo. well, Q8 of 3.6-27B. I'm building the Q8 GGUF for 3.8 now, assuming mine will finish before someone else's.
KyleJune 5 hours ago||
Others in this thread said it runs on RTX 4090.
esotericsean 4 hours ago|
Need to upgrade to a second 3090! Slowly building up my local models with Krea2, MiniMax H3 (and their new Music3), and now Qwen 3.8
More comments...