strong visual reasoning apparently, which is nice. still lacking native audio however. hoping for more companies to embrace the spirit of something like `gemma-4-12b-qat` for actual multi-modality (text, image, video, audio).
anana_ 10 hours ago||
Monstrous benchmarks! Hoping it is not benchmaxxed.
sheepscreek 5 hours ago|
I thought the same. But why claim something so shocking when it can easily be discredited and puts your reputation at risk? If they’re claiming Opus 4.6 level, I expect it to at least match Sonnet 4.6.
walrus01 3 hours ago||
Is it just me or is 3.6 27B Q8 K XL (Unsloth) holding up better in sustained token/s rate as the context fill increases over time? The token/s rate seems to be much higher for a time period deeper into context than previously seen.
At least as compared to 3.6 27B in the same quantization.
my prediction was way too far out. 4.6 at home! Woo.
pu_pe 10 hours ago||
Seems to be SOTA for its size. Hopefully independent benchmarks will come soon.
kunver 10 hours ago||
Welcome deepseek flash flash!
expedited123 10 hours ago||
Kinda was expecting to see Gemma 4 26B in benchmark comparisons :(
kamranjon 10 hours ago|
Since Qwen 3.6 27b outperforms Gemma 4 26b in most benchmarks I'm not sure the value - also Gemma 26b is a MOE model whereas this is a dense model, so not typically direct competitors at their sizes - Gemma 4 31b comparison would be interesting though.