Top
Best
New

Posted by nateb2022 18 hours ago

Kimi-K3 on HuggingFace(huggingface.co)
Related: Kimi-K3 Technical Report [pdf] - https://news.ycombinator.com/item?id=49070985
1300 points | 511 commentspage 5
spwa4 7 hours ago|
So, in normal parlance this is a 2.8T-A104B model at MXFP4 (weight) * MXFP8 (activations)

Perhaps let's call it Kimi-K3-2.8T-A104B to make matters clear.

taf2 9 hours ago||
404 is resolved let the downloads begin
colortiles 17 hours ago||
This looks really promising. Excited to see where this goes. Looking forward to trying it out!
mromanuk 12 hours ago||
Are they going to release Kimi K3.1? I’m eager to test it. According to rumors on X, it could outperform Fable. Could it be the first Chinese open-weight model to become the leading frontier model?
CodeCompost 17 hours ago||
Why is there a countdown?
broodbucket 17 hours ago||
You're not having a party?
InsideOutSanta 17 hours ago||
I think it's shameful that Moonshot isn't providing us with party kits like Microsoft did with the Windows 7 Launch Party kit. How am I supposed to properly celebrate this without fun Kimi-themed quizzes for my guests?
RALaBarge 13 hours ago||
At least they don’t make you stay till the end of the presentations to give you the software you actually came for
layer8 12 hours ago|||
It saves you from having to perform date-time calculations.
mythz 17 hours ago||
It's a release party
kansm 12 hours ago||
Wait, so I can download it and run it locally now?? Wow... But it probably won't work on my computer, right?
loudmax 12 hours ago||
The short answer is no, it won't work on your home computer. In it's current form it needs something like 594 GB of memory, far outside what you can reasonably run on normal consumer hardware in 2026.

If you have really high end hardware, you might be able to squeeze a heavily quantized version of Kimi-K3 onto your rig, but it will be too slow or too lobotomized to be useful.

This does put a near state-of-the-art open weights model within reach of what a small or medium business could afford if there's a case for local inference. It's probably not as good as Claude Fable or ChatGPT Sol. But if you're an organization that has a genuine need to run inference locally, this is a real possibility.

Is this for your homelab? Not in any practical sense.

Is this a possibility for organizations that can justify $1M or so on hardware for a near SOTA model they have full control over? Yeah, absolutely.

zozbot234 12 hours ago|||
The full K3 model will probably be way more than 594GB, that's more of a plausible range for Kimi 2.x. You'll probably be able to test run this model at full or near-full precision using SSD offload, but only at very slow speeds - probably slow enough that you'll be forced to let inferences run overnight or even spanning multiple days. Mind you, that's still useful enough for many casual users, given that they're running a near-SOTA model!
kansm 10 hours ago|||
[dead]
0x1ceb00da 11 hours ago|||
Yes it will. Buy a 4TB nvme, allocate 3TB as swap, and run a gpu emulator on your cpu.
cpburns2009 10 hours ago||
I can't imagine a GPU emulator would run better than straight CPU.
k__ 12 hours ago||
Right
hneqy2wqls 14 hours ago||
Good enough is often the right call
Aboutplants 13 hours ago|
I don’t need the fastest car to get to where I’m going. I need a car that gets me to where I’m going at the speed I’m comfortable driving at.
marvinLuck 17 hours ago||
The weightings should be released on July 27.
sreekanth850 17 hours ago||
how feasible its will be to run on modal or deepinfra? anyone here tried and tested such large models running?
thundergolfer 9 hours ago|
Modal eng here. Getting the model running is quite challenging, but its accessible right now on Modal via Endpoints: https://modal.com/blog/kimi-k3-by-moonshot-now-available-on-...
sreekanth850 7 hours ago||
15/million. that will be much higher than subscribing claude or open AI right
dsrtslnd23 15 hours ago|
maybe a quantized version on a GB300 would work? unsloth hopefully working on it.
More comments...