Top
Best
New

Posted by erdaltoprak 9 hours ago

Qwen 3.8 27B(huggingface.co)
774 points | 505 commentspage 6
maherbeg 5 hours ago|
Does anyone have a https://tenstorrent.com/hardware/cards to try it on?
g023 4 hours ago||
All this performance at such small model sizes, why are the API fees so high for the AI monopolists on this side of the world?
TomGarden 8 hours ago||
Really excited to see what people do with this. 3.7 27B was probably the best compromise between size and intelligence to run on consumer hardware
geek_at 3 hours ago|
Do you mean 3.6 27b? Because qwen 3.7 didn't have an open weight version
synergy20 8 hours ago||
I wish this can run directly on my RTX 4090, seems like 30B is the sweet spot for dense model to run locally, sadly RTX 5090 is very expensive and I need a new PC and new power supply(and UPS) to run that, adding a second RTX 4090 is another option, but not sure if my PC can do that yet.
baron3dl 8 hours ago||
even a 3090 will give you the VRAM headroom. i run Q8 on an 3090/A6500 combo. well, Q8 of 3.6-27B. I'm building the Q8 GGUF for 3.8 now, assuming mine will finish before someone else's.
KyleJune 6 hours ago||
Others in this thread said it runs on RTX 4090.
kunver 8 hours ago||
Looks like a pretty significant improvement on the DeepSWE benchmark compared to the previous 27B model.
bertili 7 hours ago||
Wow. Speed improved as well. 200t/s on a RTX 5090!

https://x.com/sgl_project/status/2088281320422322413

kristopolous 8 hours ago||
q4km is about 48 tps on a 4090. my llama.cpp params are --flash-attn on --parallel 1 --load-mode mmap
m_ke 8 hours ago|
With spec decode should easily get to >100tps

on my dual 3090s qwen 3.5 27b was running at around 110tps using the config from https://github.com/noonghunna/club-3090

make that 200tps on a single 5090, 4x faster than opus https://x.com/radixark/status/2088285681131110446

devs about to get handed a two 5090 box each and told to max that out

esotericsean 6 hours ago|
Need to upgrade to a second 3090! Slowly building up my local models with Krea2, MiniMax H3 (and their new Music3), and now Qwen 3.8
More comments...