H
Hacker News
Top
Best
New
Posted by mft_ 10 hours ago
Flash-MoE: Running a 397B Parameter Model on a Laptop
(github.com)
262 points
|
92 comments
page 3
pdyc 9 hours ago
|
next
[-]
impressive, i wish someone takes a stab at using this technique on mobile gpu's even if it does not use storage it would still be a win. I am running llama.cpp on adreno 830 with oepncl and i am getting pathetic 2-3t/s for output tokens
qcautomation 2 hours ago
|
prev
|
next
[-]
[dead]
Yanko_11 4 hours ago
|
prev
|
next
[-]
[dead]
leontloveless 4 hours ago
|
prev
|
next
[-]
[dead]
robutsume 6 hours ago
|
prev
|
next
[-]
[dead]
aplomb1026 4 hours ago
|
prev
|
next
[-]
[dead]
leontloveless 6 hours ago
|
prev
|
next
[-]
[dead]
diablevv 7 hours ago
|
prev
|
next
[-]
[dead]
leontloveless 8 hours ago
|
prev
|
next
[-]
[dead]
jee599 7 hours ago
|
prev
[-]
[dead]
More comments...