Posted by bastitx 14 hours ago
additional paper: https://tej.as/blog/aleph-alpha-kolibri
How is a 27B dense model bigger than a 78B MoE?
Congrats to the team.
If they can follow Qwen3.8-Flash-Next, the could draft off the huge reduction in VRAM requirements.