All this performance at such small model sizes, why are the API fees so high for the AI monopolists on this side of the world?
TomGarden 8 hours ago||
Really excited to see what people do with this. 3.7 27B was probably the best compromise between size and intelligence to run on consumer hardware
geek_at 3 hours ago|
Do you mean 3.6 27b? Because qwen 3.7 didn't have an open weight version
synergy20 8 hours ago||
I wish this can run directly on my RTX 4090, seems like 30B is the sweet spot for dense model to run locally, sadly RTX 5090 is very expensive and I need a new PC and new power supply(and UPS) to run that, adding a second RTX 4090 is another option, but not sure if my PC can do that yet.
baron3dl 8 hours ago||
even a 3090 will give you the VRAM headroom. i run Q8 on an 3090/A6500 combo. well, Q8 of 3.6-27B. I'm building the Q8 GGUF for 3.8 now, assuming mine will finish before someone else's.
KyleJune 6 hours ago||
Others in this thread said it runs on RTX 4090.
kunver 8 hours ago||
Looks like a pretty significant improvement on the DeepSWE benchmark compared to the previous 27B model.
bertili 7 hours ago||
Wow. Speed improved as well. 200t/s on a RTX 5090!