Top
Best
New

Posted by bastitx 13 hours ago

Kolibri: A Sovereign Open-Weight Model(aleph-alpha.com)
tech report: https://aleph-alpha.com/downloads/tech-report.pdf

additional paper: https://tej.as/blog/aleph-alpha-kolibri

467 points | 280 commentspage 3
dosinga 7 hours ago|
> The second was to rephrase German documents we already had. An LLM rewrites an organic German document in the style of an encyclopedia entry, a Q&A dialogue or a text passage, preserving its content.

"an LLM" -- does that mean they are effectively learning from that LLM the German encyclopedic style? makes me wonder which LLM and how that is really sovereign.

UncleOxidant 3 hours ago||
78B MoE with A3.6B is a very nice size.
petesergeant 9 hours ago||
I wish nothing but luck for an EU model, but:

> intellectual-property safety

My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.

rpdillon 9 hours ago||
This doesn't seem to be true. There's a clear legal path via the first-sale doctrine to train models on copyrighted works. It's been years now, and publishers still don't seem to be offering anything for training (e.g. bulk licenses solely for training use), but adversarial interoperability via cutting up books and scanning them remains perfectly legal.

There's also the ability to distill other models, which is also not illegal (though I'm sure they like to come after whomever for TOS violations, but thats a civil matter).

And, of course, the obligatory copying-isn't-theft observation. A recent supreme court judgment put it well.

> Since the statutorily defined property rights of a copyright holder have a character distinct from the possessory interest of the owner of simple “goods, wares, [or] merchandise,” interference with copyright does not easily equate with theft, conversion, or fraud. The infringer of a copyright does not assume physical control over the copyright, nor wholly deprive its owner of its use. Infringement implicates a more complex set of property interests than does run-of-the-mill theft, conversion, or fraud.

Folks are pretty smart here, I think we can handle these nuances, even if we don't agree about whether they are good.

Edit: reading through the full text of their post, it looks like they are using common crawl, which is likely just as much of a copyright infringement as Anna's Archive -- it's not like published works have a unique claim to copyright. I think this strengthens your point, though: I was expecting to see scans as training data, but it doesn't appear to be the case.

mapontosevenths 7 hours ago||
"The Congress shall have Power To ... promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries." - The United States Constitution

Copyright is a government mandated monopoly that was only granted in order to advance the arts and science. Any interpretation that runs contrary to that is bollocks being used by the religiously or financially motivated to serve their own petty interests to the detriment of societies.

befelix 5 hours ago|||
Disclaimer: I am part of the team that trained Kolibri, opinions are mine.

You're right that especially big models benefit from training on copyrighted material in terms of world knowledge (especially from books). However, in the small model space imho agentic capabilities where the model looks up knowledge on the fly are much more important. That's what we focused on quite a bit during training. Personally, I also don't think stealing stuff is okay.

petesergeant 5 hours ago||
I hope you're right, but I guess we'll see when we get independent benchmarks. My intuition is that even for small models and models mostly focused on tool calling, you _still_ need all that extra contextual stuff for the magic, but I am further from the coalface than you are.
jamienk 4 hours ago|||
Why should we concede this "stealing" framing? If I want to put texts into my computer program, why should that count as copyright violation? I think that idea is as ridiculous as saying that reading a document is a copyright violation.

Regulate large cloud services and proprietary software - yes! But not on the basis of "Intellectual Property".

The "legal" issues here are very very complex and we should not passively wait for or accept corrupt court rulings, international trade agreements, proposed laws, or worst of all propaganda that pushes a parochial and craven view on this.

ekidd 9 hours ago|||
This model isn't terrible, at least on the benchmarks. It's 78B A3B and performs about like Qwen3.6 35B A3B. You can probably run it comfortably in 96B of RAM with a decent quant that doesn't lose too much.

Unfortunately, Qwen3.6 35B A3B isn't really a useful coding model. You'd probably want Qwen3.8 27B at a minimum, which requires at least 32GB of VRAM (not system RAM) to run semi-comfortably.

So this isn't going to be a competitive model for hobbyists, and you'd have to be a bit desperate to use it for coding. But if you work in a regulated industry and don't mind paying for a bit of extra hardware, it isn't catastrophically bad, either. Probably would work fine for information extraction or as a "classifier" like Jev. (Almost any GGUF model can be turned into a classifier using llama-server. See pi.dev codemode for sample code.)

So they're not a real contender yet, but they look like they're probably at least minimally credible.

torginus 8 hours ago|||
hasn't IP law passed the statute of limitations? As in most models are probably trained on output of other models, as creating enough data otherwise is not feasible. Additionally, they are trained on github repos made since the AI boom, which were generated by models with IP issues (who knows what and how).

Thus training on 'clean' data is like trying to unscramble an egg.

miohtama 7 hours ago|||
There is no word for copyright in Mandarin :)
dotancohen 6 hours ago|||
Is this true? Do they possibly use a loanword or a descriptive term? Certainly you are not implying that the concept of copyright does not actually exist in Chinese society?

For what it's worth, in my language we don't have a word for copyright either. We have the concept, though, we just call it literally Creators Rights זכויות יוצרים and the borders of what is and what isn't covered broadly map to the familiar concepts of IP.

applicative 3 hours ago||||
版权
cheez-wiz 6 hours ago|||
I mean, there is one, they have copyright law. Forgive me for being slow is this a joke about the widespread theft of IP in China? Or was the acquisition of training data just much more 'accepted' in China compared to the west?

I feel I messed up your quip =/ I'm new here, go ez. Not looking for excuses to hate on China either.

miohtama 1 hour ago||
Yes, the law exists, and anyone can go and ask their IP back in the court of Beijing.
andy99 7 hours ago|||
Not really. The upside of competitive newer models is all in the proprietary data they are trained on. This is why data labeling, RL environments et al have been such a big industry, OpenAI and Antrhopic are paying literally billions to get the data they need. Do people think the ability to do research level math or advanced cybersec comes from just training on more public data?

A real sovereign effort could invest heavily in this, whatever people accuse China of “stealing” I’m sure they are also generating tons of their own data and are probably the primary sovereign doing so outside the US labs.

embedding-shape 9 hours ago|||
Have you tried the model itself and seen if it's "even slightly competitive" or not, and have specific complaints about it? Otherwise it feels like you're complaining about something that is easy to test but rather than taking the time to actually figuring that out first, you're arguing about some general and theoretical thing which the submission (may) directly disprove.
petesergeant 8 hours ago||
No, I haven’t, but I’ll donate $20 to the non-political charity of your choice if it doesn’t turn out to sit a significant difference from the frontier.

I think it’s a safe assumption that they’re leaning into “sovereign” because performance is bad.

ygjb 8 hours ago||
I think you have it backwards. Sovereign is the goal, good can come later.

There is a proliferation of sovereign models under development specifically to address data sovereignty, and a loss of performance is absolutely acceptable over the risk that a once ally will turn adversarial, or a foreign business stops serving what has become critical infrastructure.

Zambyte 9 hours ago||
Is it noble? The entire notion that training data can be "stolen" at all is quite silly. If I "steal" content that someone created to use for training, what am I actually stealing? They didn't lose anything. They still have everything they had before. What was "stolen" was "unrealized profit", or put another way: money that wasn't theirs, that they had no entitlement to. The only actual crime that is committed is "unauthorized copying", not stealing. Support and enforcement of copyright feels wildly authoritarian. It's hard to see it as noble.
folkrav 8 hours ago|||
The same could be said of any digital product being sold. Nobody actually loses anything but the actual sale either when you download a cracked game or piece of software, a movie, music, etc.
pepperoni_pizza 8 hours ago|||
That's fair, but then people like you complain when someone "steals" I mean distills openai or anthropic models.
brookst 8 hours ago|||
It’s poor form to argue against someone by imagining something totally different that they might believe, which would then make them hypocritical.
Zambyte 8 hours ago|||
I don't complain about that. Model distillation is excellent.
lmf4lol 1 hour ago||
Wow this is so cool. Glad that Aleph Alpha does that after Mistral threw the towel in the ring (and disappointed the european AI crowd massively!!!!). After AA got sold to the Canadians, I thought its over but this is a really cool comeback and the depth of the tech report shows that they a serious about openness. I hope I can use their model soon in my product. Would be awesome to have a European model to offer!!!! I really wonder how it compares to deepseek v4.1 flash
wg0 4 hours ago||
Anyone thinking this won't improve or is behind etc is blatantly wrong. It'll catchup within a year. Like that unknown wise and visionary man inside Google once said about their competitors: We have no moat neither does anyone else."

Congrats to the team.

peterBlue75 3 hours ago|
Thanks! What matters to us right now is not so much our current position, but the velocity with which we’re moving. Kolibri 1 is the first measurement of our position, there’ll be more. We want to build this in Europe and this is only the start. The report also shows our focus on building out proper tooling with our Model Factory.
tosh 9 hours ago||
i wonder if the custom tokenizer is better in practice, the examples look interesting though
Wittie 3 hours ago|
[dead]
Jeeetendra 7 hours ago||
3.5b active params sounds cheap until you remember all 78b still has to fit in memory. curious what the smallest practical self-hosted setup looks like for german docs.
layer8 7 hours ago|
The article talks about what setup is needed.
pythonic_hell 10 hours ago||
The benchmarks are impressive given the problem space they are working in.
rrr_oh_man 9 hours ago|
German government?
sajithdilshan 9 hours ago||
More like Chancellor’s Office
d2kx 10 hours ago||
German here. We are cheering for Mistral, which is making some good moves before the year is over, and Black Forest Labs for non-coding. But that's about it.
hypfer 9 hours ago|
Speak for yourself and yourself only.
CorezIoOfficial 9 hours ago|
Im surprised by how well this works. What is the difference from this and union alpha (other than the fact that it is open weights)?
More comments...