Top
Best
New

Posted by marcobambini 7 hours ago

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s(github.com)
84 points | 33 commentspage 2
logicallee 3 hours ago|
Interesting project. The headline number (29 GB of RAM) is for 4k context.

From what I've read elsewhere, Kimi K3 is quite verbose in its thinking. At the quoted rate, it would generate only a total of 1.8k tokens in 1 hour. Is that enough for it to get any thinking done and produce output on more complicated prompts?

0cf8612b2e1e 2 hours ago|
I saw someone’s excellent idea that if you have a slow system like this, you should communicate by email. It is no longer meant for realtime iteration, but more pointed questions for which there is more effort and time expected on both parties.
rwz 1 hour ago||
0.5t/s is still too slow even for email. For a moderately large inquiry (1MTok output, let's ignore the 4k context window limitation for now) it'll take the model around 23 days or uninterrupted execution to answer a single email.

Real world inquiries are gonna be much slower of course, but this setup is still too slow to do anything meaningfully useful I think.