Top
Best
New

Posted by bastitx 13 hours ago

Kolibri: A Sovereign Open-Weight Model(aleph-alpha.com)
tech report: https://aleph-alpha.com/downloads/tech-report.pdf

additional paper: https://tej.as/blog/aleph-alpha-kolibri

444 points | 274 commentspage 2
tornikeo 7 hours ago|
I love how the mere mention of a "sovereign" in LLM's announcement is the declaration of defeat.

This thing is worse than a Qwen3.8 27B.

tfburns 3 hours ago||
Hi there! I'm Tom. I worked on Kolibri at Aleph Alpha :)

I think it's not simple to directly compare a 27B dense model with an MoE model like ours. As we know, dense models need all params active for every token. Whereas, MoE models (especially sparse ones like Kolibri) fewer active parameters and correspondingly less compute per token.

Among the MoE models we compared against in our tech report and model card, though, Kolibri performs very well in our evaluation, including against models with 12B active parameters. It also best model in the group we tested within that range of active params.

So, I think it's fairer to see this as a trade-off. Kolibri needs less compute per token but more memory, while Qwen3.8 27B needs far less memory and more compute per token. In the report, both are actually on the quality-vs-serving-cost Pareto frontier among the models we evaluated, just at different points.

andy99 6 hours ago|||
It’s an incentive problem. If “sovereign” becomes your claimed value proposition, you can claim success even if the models not competitive, so nobody is pushed sufficiently hard to actually make it good.

Sovereign works when talking about building a commodity supply or something, not in literally the world’s most competitive and fast moving field.

Those seeking sovereign capability would be better off aiming to be best at something, even something much narrower than an all round LLM. Or just fast following and making something that matches leading performance, which is close to what the Chinese labs do currently.

okamiueru 5 hours ago||
I can think of at least one other pretty good reason, which is in anticipation of regulatory capture. If "LLM used must be FOOBAR-certified" and coincidentally no Chinese models can get this certification, having such an alternative is a lot more valuable than just scoring highest in a set of benchmarks. Not to mention that these benchmarks aren't always accurate.
glitchc 6 hours ago||
It doesn't help that Qwen3.8 27B is an excellent model.
gpugreg 1 hour ago||
I was wondering whether the model could help me with German bureaucracy. Unfortunately, the answer is "no".

More specifically, I asked the model about what I should put on my contact page, which can cost you in the order of 500 € in Germany if you don't write the right magic words.

Kolibri incorrectly referenced the "Telemediengesetz" ("telecommunication act"), which has been superseded by the "Digitale-Dienste-Gesetz" (DDG, "digital services act") since 2024. The model knows about the DDG, but does not reference it unless specifically instructed to do so.

If anyone of the developers reads this, you can fix this by introducing a recency bias during training. You can even control it by conditioning the model on a date provided with the system prompt or first prompt, so you can travel in time.

mark_l_watson 6 hours ago||
Looks interesting. I just went to download from HF, but they only have fp16 which won't fit on my Mac.

Good to see Europe adding toe what Mistral is doing. +100

befelix 3 hours ago|
FP8 weights are actually the default and what we trained for: http://huggingface.co/Aleph-Alpha/Kolibri-1

We have a separate repo for BF16: Kolibri-1-BF16. Will still be tough to put it on a Mac though :(

Disclaimer: I'm part of the team that trained Kolirbi

vzaliva 7 hours ago||
"Languages: German and English" – this is odd. That means their dataset is limited. In my understanding, frontier models are trained on multilingual datasets and can combine knowledge no matter what language it was written in.
AnonymousPlanet 3 hours ago|
Have you read the non-English output of those mutlilingual models? They say they put extra effort in to improve nuances and tone in this one. That alone is worth having in many settings.
peterBlue75 2 hours ago||
Related thread on the tooling:

Model Training as Code - https://news.ycombinator.com/item?id=48673450 - June 2026 (24 comments)

satvikpendem 2 hours ago|
And related thread on the actual model post: https://news.ycombinator.com/item?id=49942706
thatguysaguy 3 hours ago||
> 3B active

> built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace

Cool to be doing more independent model lineages, but I sure hope no one actually uses this as part of any aerospace engineering...

james45 7 hours ago||
the sovereignty topic needs more attention in general so great to see. self-hosting the model is one piece of sovereignty, but how do we handle the rest of the agent stack - embeddings, retrieval, memory, etc. Has anyone put together a practical agent stack that's 100% sovereign, where they control it all?
kouunji 4 hours ago||
If you mean hosting and serving all the other elements of the agent stack, then, yes. We handle sovereign data, so part of our whole value proposition is that every part of our pipeline is hosted in Canada.
magus-stoopr 6 hours ago||
I'm looking for this too
lmf4lol 1 hour ago||
Wow this is so cool. Glad that Aleph Alpha does that after Mistral threw the towel in the ring (and disappointed the european AI crowd massively!!!!). After AA got sold to the Canadians, I thought its over but this is a really cool comeback and the depth of the tech report shows that they a serious about openness. I hope I can use their model soon in my product. Would be awesome to have a European model to offer!!!! I really wonder how it compares to deepseek v4.1 flash
cheesecakegood 6 hours ago||
For those sick of “Pareto frontier” talk, just shorthand it as “it’s the best at some very particular thing”. Obviously that one thing/tradeoff it’s good at may not necessarily be compelling, but it is either a loose sign of quality, or a sign that they’ve chased some tiny edge into the ground.

I’ll be curious to see which it becomes in the next year - nba “very narrow record”, or a sign you can hang with the big boys.

UncleOxidant 3 hours ago|
78B MoE with A3.6B is a very nice size.
More comments...