Posted by snehesht 10 hours ago
That's a neat number (576 is the square of 24). Ofcourse it must have come from 24 * 2^10.
Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.
So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.
FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.
One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.
Progress on running local models has been amazing.
So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2
https://huggingface.co/nathansutton/Qwen3.8-27B-Ternary-Bons...
or a MoE retrofit like Qwen3.8-35B-A3B with or without mtp
https://huggingface.co/NovaeonStudio/Qwen3.8-35B-A3B-Distill...
https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...
It is more useful than Qwen3.8:27b (which is already quite good) and runs faster on my 7900 XTX / 64 GB DDR4 system.
Local LLM is getting more exciting every day!
amazing project, congrats on the launch
That's the exciting part of this - before the best you could run on <24gb vram was qwen3.8-27b at q4 quantization. Now you can run a nerfed 125B parameter model on under $800 of hardware, and it beats a less-nerfed 27b model.
The Readme doesn't say, but it's all AI generated, so..