Awesome results. Opens up doors for a lot of people.
abraxas 13 hours ago||
I'm not following the local mdoel scene too closely but this seems quite amazing. Is this able to be run on Apple silicon too?
Havoc 13 hours ago||
Their first 27B bonsai was able to run on an iphone.
kamranjon 13 hours ago||
"Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient than an 8B model running in full-precision."
pizza234 13 hours ago|||
Their mention of the 5090 is bit odd, since on 32 GB GPUs, Q6 fits while having better quality. Very interesting model for 16 GB GPUs though!
pwython 6 hours ago|||
Sometimes you want a decent model running in the background that doesn't take up all the VRAM.
blurbleblurble 6 hours ago||
Or maybe even to run parallel threads of the same model!
sisve 12 hours ago|||
They mention 5090 with regards to speed, Q6 will not have that speed?
And speed matters a lot for many use cases
selectodude 12 hours ago|||
150 tokens per second on a ternary model implies that it’s GPU bound, I’d bet a Q6 model is even faster because it’s existed longer and seen more optimization. You’d have to be insane to not run an NVFP4 quant over a ternary quant on Blackwell if they both fit.
wincy 6 hours ago|||
With Ninfer and Qwen 3.8 27b it uses a groupwise int mixed quant, and it gets 160 tokens/sec. The mixed quant is between 4 and 6 bits.
azatom 12 hours ago|||
it is like "my fridge is 2mkm (millikilometer) from my desk"
m=0.001 h=3600 it should be just Ws or just J
(In case anyone remembers the compression post from yesterday[1], I checked and this one doesn't qualify for further compression - it's not zero-biased at all.)