Top
Best
New

Posted by erdaltoprak 14 hours ago

Qwen 3.8 27B(huggingface.co)
976 points | 627 commentspage 11
dude3 10 hours ago|
[flagged]
fintuner 13 hours ago||
[flagged]
RobertasTa 13 hours ago||
[flagged]
steffi_oliver 12 hours ago||
[flagged]
alpha_trion 14 hours ago||
NICE, i've been waiting for this drop, thanks for posting this
literoldolphin 11 hours ago|
Why is anyone even using video cards these days? You may as well be burning cash.

This is the perfect candidate for just splattering it on your nvme and then reading it off there and into memory. All of these run perfectly fine on simple m4 silicone:

https://github.com/drumih/turbo-fieldfare

https://github.com/leonickson1/Swiftlet

https://github.com/sqliteai/warp

awkwardpotato 11 hours ago||
Those are all for MoE models. And I prefer measuring my tokens in t/s instead of s/t
ferrouswheel 11 hours ago||
Lol, "burn money on apple hardware instead!"
literoldolphin 11 hours ago||
And yet it's also a laptop you can basically take anywhere unlike a giant video card with 1000 watt power supply requirements.
bigyabai 6 hours ago||
Nvidia and AMD both have their own unified memory laptop SOCs, now. Apple Silicon's GPU is relatively weak, it's one of the less-efficient ways to use 100w for compute.

Even the fastest Apple Silicon chips like the M5 Max and the M3 Ultra still put up worse GPU compute performance than last-gen laptop RTX 4080 chips. And they don't scale, the largest M3 Ultra cluster you can configure is still ~2,000x smaller than a DGX SuperPOD. There's a reason Apple discontinued their rackmount hardware, there's very little demand for Apple Silicon in the datacenter.

literoldolphin 15 minutes ago||
As I mentioned above you can't take the data center into the cafe somewhere. We're talking about running local models here.