Top
Best
New

Posted by bastitx 12 hours ago

Kolibri: A Sovereign Open-Weight Model(aleph-alpha.com)
tech report: https://aleph-alpha.com/downloads/tech-report.pdf

additional paper: https://tej.as/blog/aleph-alpha-kolibri

444 points | 274 comments
miellaby 6 hours ago|
The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
ivo-42 4 hours ago||
I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
brcmthrowaway 4 hours ago||
How do you cleanse the data at this scale?
ivo-42 4 hours ago||
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.

We have a lot of details in the tech report if you want to go deeper.

stephantul 2 hours ago|||
Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.

I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!

xvfLJfx9 1 hour ago|||
Are there plans to make much larger versions of this model? With 500B-1T params for general purpose knowledge tasks, similar to the current leading proprietary models?
idiotsecant 1 hour ago||
How exactly would an open project do that?
zelphirkalt 6 hours ago|||
In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.
davidjfelix 50 minutes ago|||
to be fair, this discussion has been had numerous times here and the industry has arrived on "open-weight" to describe the practice of releasing the post-training weights in an open manner but not releasing the data it was trained on.

That's what they call this and I think it's a pretty clear definition these days to people in the industry.

slow_typist 57 minutes ago|||
How, is the training data public?
amelius 1 hour ago|||
I wish universities would take it upon themselves to curate the training sets for these models.
dang 3 hours ago|||
The paper mentioned is here: https://tej.as/blog/aleph-alpha-kolibri.

(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we're merging the threads.)

p-e-w 6 hours ago|||
This alone makes it much more valuable than many high-profile releases despite not quite performing at the same level.
kingcauchy 5 hours ago|||
Yeah the pdf alone is awesome as a learning tool.
zwaps 3 hours ago|||
Such a crazy change from the times of Luminous, when they published a three pager with a claim that the model is similar good as „gpt 3“ (which??) with some graphs without y axis.

Bravo team!

ofjcihen 6 hours ago|||
Hopefully this becomes the new standard.

It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.

Loquebantur 5 hours ago||
Be cautious what you wish for. Tools don't tell you what to do with them.

Open source LLMs "democratize" access to the "intelligence booster" that is AI. But while that has several benefits, it also has several downsides.

Humanity has the serious problem of being underdeveloped in the "spiritual" department. Ethics is often considered some sort of lifestyle choice, but it's actually the difference between order and chaos in a society.

Everybody being able to do anything means somebody will be able to do something you don't like. At an arbitrary scale.

adrianN 4 hours ago|||
Of course openness is only worse than leaving everything under the control of a select cabal of you believe that cabal to be more ethical than the rest of us.
zanderwohl 4 hours ago|||
The select cabal who believe in an eschaton they're actively trying to bring about, as well.
skinfaxi 1 hour ago|||
Would you apply this reasoning to the proliferation of nuclear weapons?

edit: why is the parent rationale sensible for AI and not nuclear technology?

orbital-decay 58 minutes ago||
Absurd and incomparable.
skinfaxi 49 minutes ago||
Why? That seems like a shallow dismissal. Why is a world changing technology okay in the hands of a small cabal in the one case and not another?
orbital-decay 41 minutes ago||
"Why is the wheel not okay to gatekeep but the nuclear bomb is?" Even the framing is manipulative from the beginning. By asking this question you're already assuming they're in any way comparable. They are not even remotely equivalent and the entire comparison is utterly absurd.
skinfaxi 30 minutes ago||
AI is more like a nuke than a wheel. Do you disagree? Can you suggest a less manipulative framing? I am personally a proponent of open source AI and models but I found this cabal framing strange when we do indeed rely on this kind of control for other world-altering technologies. And AI is different in that it enables technological development in ways quite unlike the wheel in a general sense.
Zigurd 5 hours ago||||
The main thing that bugs me about the risks discussion around AI is the lack of specificity. Commenters here have a good grasp of the risks around finding vulnerabilities faster than they can be patched. That's good and it matches the applicability of LLMs to coding.

But the applicability and the ROI of LLMs for other use cases than coding is a lot squishier. Also correspondingly the risks are unspecific.

As for what to do, ethical disclosure of vulnerabilities provided a good framework for disclosing software vulnerabilities discovered with the assistance of LLMs. What is going to be novel and calls for our spiritual development in other domains?

GolDDranks 4 hours ago|||
I find it odd that more people don't realize that we are talking about risk of elevated *general intellectual-domain capabilities* as a resource. To be clear, I'm not claiming that LLM + RF is necessarily THE technology that poses the risk, the risk is in recursive self-improvement and whatever technologies will result.

It's very clear to that any specifics couldn't capture the risks, because the capabilities, including the risks, are one level higher than any specific techonolgy. It is the process of advancing technology itself, in accelerating speed, that poses the risk.

Zigurd 3 hours ago|||
The reason I use coding and vulnerabilities as an example is that it is a concrete example. It is what people pay for now when they buy AI. And the risks are specific and can be examined in detail. Some threads on this board currently show that even these more concrete and specific risks are often overblown, with LLMs finding low risk bugs and sucking up resources to evaluate and fix them.

Here you are claiming that AI products are going to reach AGI or RSI in the foreseeable future. Of course you can't "capture the risks" with specifics because those are inherently unspecific futures. It's a bit like saying when we invent antigravity all hell will break loose.

I would believe those future risks more if there were a progression of risks. What other than finding vulns has those characteristics?

Loquebantur 4 hours ago|||
What poses the risk is the combination of abilities past a certain point enabling you to do basically anything.

While being unable to judge whether you should in the first place.

vincnetas 2 hours ago||
ai cant move atoms at unlimited rate and also have limited energy. so your claim that "basically anything" is a bit of a stretch.
Loquebantur 4 hours ago|||
You're right, people weirdly lack imagination on what "higher intelligence" (minus ethics) actually affords you, let's have a look:

What do average people currently want? They're taught, the most important thing was being rich. So they will ask their AI to make them rich. Most real life ways to get there are "sketchy" to say the least, usually downright unethical and anti-social, but US society turns a blind eye when the "Wolf of Wall Street" comes out on top and the schemes don't easily fit into average people's abilities of moral judgement.

-> Large parts of US society suddenly engaging in all kinds of "semi-legal/hyper-illegal" fraud schemes, at the expense of already saturated environmental and societal resilience. Guaranteed collapse.

Or, let's get rid of those pesky neighbors/wrong-colored people/annoying opinions? Again, "legal" is a pretty squishy concept and only really applies when you don't have the legal expertise to get around it. Now you can.

Or, look at the basics: what is "real"? You only "know" because you trust certain people and institutions. Generative AI can help with that /s.

It's not only about "building weapons of mass destruction". It's about doing the same shit as usual, but a thousand times faster/amplified. Look up poly-/metacrisis for starters. Going faster with AI when there's a wall in front of you isn't the best idea.

Zigurd 3 hours ago|||
This is analogous to the problem of spam, which is a problem about five minutes younger than email. Before spam you had to buy ads in the back pages of magazines you think target vulnerable demographics. Meta already spews fraudulent ads in horrific volume.

In other words, ambitious frauds have already explored all of the angles and bought all the ads. At worst, LLMs will create a few more successful but less ingenious frauds.

fwn 3 hours ago|||
Whenever classic p(doom) sentiments are explained through a text with obvious LLM markers, I wonder whether I am looking at a superhuman persuasion attempt.

..or maybe the commenter did look at superhuman persuasion long enough to believe it would be best to channel those ever the same fear fantasies from the LLM through their account to the reader.

On a more serious note, just look at the doom premises here: "Large parts of US society suddenly going criminal" is from the movie "The Purge", I think. It is fiction.

The idea that generative AI takes away our ability to find out reality. ... I don't know. People write about that a lot, but it still seems very far fetched.

Maybe through some terminally online overconsumption, like with social media? I wouldn't know.

With new AI capabilities we will have to adjust, I am sure. Media, science, education and law are changing very visibly right now. Those p(doom) narrations just seem to be pre-IPO hype though.

It is just so so dangerous. That is why they want to go public and only want to care about optimizing for the next quarter ...right before breaking into AGI. /s

computerdork 5 hours ago||||
Ah, didn't think of this. Some rogue militia group might try to use this LLM (or create their own LLM based on this work) to help them create biological weapons or to do a mass hacking the infrastructure of targeted country.

Wonder what safeguards Kolibri uses to prevent this? Or if they even can

zanderwohl 4 hours ago|||
I don't think that AI is as much of a boost to bioweapons as people think. Lab work doesn't get easier just because the experiment design part does.
Zigurd 5 hours ago|||
Chemical and biological warfare is hard. A cult in Japan created a mass casualty event using nerve gas. Which is the only somewhat "successful" terrorist WMD attack I know of. Knowledge of how to create these weapons isn't new. Guns and bombs are the most widely used terror weapons for a reason.
UberFly 4 hours ago||
Enter LLMs to help through all those pesky hard parts.
shawabawa3 4 hours ago|||
The hard part is probably finding the lab equipment and chemicals without being noticed, and choosing not to use it to manufacture drugs instead which would be much more profitable
throwaway27448 3 hours ago|||
Intelligence was never the bottleneck tho
happosai 1 hour ago||
For terrorists it is tho. Four lions is basically a documentary. The only clever and creative terror attack happened 25 years ago.
skinfaxi 1 hour ago|||
Didn't Mexican cartels kidnap telco workers to build them separate infra? We expect terrorists to be less resourceful?
throwaway27448 1 hour ago|||
What's stopping them from using claude today? What could anyone possibly do to stop them from using open models? This line of thought seems like corporate/political/pr pandering more than a meaningful concern.
ofjcihen 4 hours ago||||
Oh definitely, and I’m in the cybersecurity space so I’m already on the “worst case scenario committee” hah.

But the alternative just seems… so much worse to me?

A select few groups gating access to the ability to do everything seems like neo-fuedalism in the making.

And to be fair even the gating that we do have (daybreak, CVP, etc.) is already being circumvented via keys being stolen and sold on the dark web.

Loquebantur 4 hours ago||
Yes, a "select" (rather, self-selected) elite "controlling" AI according to their wishes, what could go wrong?

Clearly not a "better" scenario. The real problem though seems people feigning helplessness? You can't leave society "to its own". You are part of it and go where it goes. So better start steering.

When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that. Just like you shouldn't be allowed to drive a car or fly a plane or command a rocket without proper guardrails, safeguards, prerequisites, etc.

"General" intelligence isn't present in humans, why does it need to be in AI?

skinfaxi 1 hour ago|||
> When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that.

What do you mean "cannot"? As in you are granted abilities that have no responsible use?

throwaway27448 3 hours ago|||
> When access to AI gives you abilities you cannot use responsibly

I don't think this is realistically a problem at all. It just makes certain types of research cheaper and less time-consuming. And, again, this is also a problem with the american services.

throwaway27448 3 hours ago||||
I think the unethical things are happening already with boutique firms. I admit I don't get the concern.
OtomotO 3 hours ago||||
> Humanity has the serious problem of being underdeveloped in the "spiritual" department.

As is shown to us by filthy rich people every day.

Or did you mean the burglar in the fawellas?

98753579909754 4 hours ago|||
[dead]
api 8 hours ago||
His point about regulation and innovation is great and I wish more people thought like that.

One of humanity’s biggest problems here is we don’t know how to do moderation.

We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.

The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.

I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.

dang 3 hours ago||
(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we've since merged the threads, so I've moved it into the subthread which is specifically about the paper being responded to.)
andai 4 hours ago||
>We trained Kolibri with abstention data and with our Merlin-Arthur protocol. As a result, it is trained to say "I don't know" when the answer isn't in the context.

https://aleph-alpha.com/en/blog/bounding-hallucinations-merl...

Glitch752 55 minutes ago||
I'm not super impressed, at least in English. I asked:

  What is the canonical interpretation of the song ‘Glass Flowers’ by Armand Uso?
(A made-up song and name) And received:

  "Glass Flowers" by Armand Uso, a track from the 1978 album The Art of Falling in Love, is generally understood as a melodic reflection on the fragility and impermanence of love. [...]
I saw similar responses for other questions. Qwen3.8-27B correctly refused without web search.
Velocifyer 3 hours ago|||
Just like Merl.

https://minecraft.wiki/w/Minecraft_Support_Virtual_Agent#Sys...

(for context, the Minecraft Support chatbot, aka Merl, had a meme because she kept saying "I don't know" to questions like "How to craft a diamond pickaxe".

swozey 4 hours ago||
Gemini gives me a lot of hard-no responses, and I think to myself, "well can you find out? what am i paying you for" usually

Gemini: Here's a list of links to check

kkm 7 hours ago||
Thank you Aleph Alpha team for making it open.

We as many other’s were curious to try and benchmark it.

On that note, as a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few days.

No GPU. No setup. Just try it. tesseracted.com/kolibri-1-chat/

https://x.com/konarkmodi/status/2106373678589960260?s=46

trvz 6 hours ago||
Friendly note: your website's font at its current size is pretty bad on a non-retina display before zooming in.
moritzwarhier 4 hours ago|||
It's been surprisingly good at the things I threw at it (history, culture and conversing, web search)!
kkm 3 hours ago||
I agree.
tharkun__ 6 hours ago||
Not impressed. I asked it how to run itself (giving it the Huggingface link) on limited RAM i.e. less than stated as needed and on llama.cpp and true to what we read about "it will tell you when it doesn't know" that's almost all I got: It doesn't know, it told me I should go click on tabs in the Huggingface interface for more information. This was with extended thinking on.

No, I'm not gonna do that, I asked you to do that Mr Kolibri.

Also feedback on that interface: It's very annoying while answering. It almost immediately shows a list of sources, which on my screen fill up all the space and then when it starts answering it keeps those in view but also scrolls down the tiny part of actual text its outputting but I can't scroll up to start reading from the top, coz it keeps scrolling. I have to wait until it's completely done generating its output.

sigmar 4 hours ago|||
It can do search, but it doesn't seem to have a tool that pulls URLs into context. I get why you would expect that tho, as most productized LLMs do it.
kkm 3 hours ago||
Correct, search is added as a tool.

Uses Brave search, the idea is to test how well the Model can decided when to leverage search or not.

tharkun__ 2 hours ago||
And that's my point when I said I wasn't impressed: It apparently can't do that properly. It can't seem to think things through by itself and then make more tool calls to fetch more information.

What I couldn't tell but maybe you can tell us: it also complained that it couldn't read the full text as something was cut off. Is that because of the tool you gave it, of brave itself or is it the model?

kkm 5 hours ago|||
Thank you for the feedback on UI, improved the streaming to make it less frustrating.
peterBlue75 4 hours ago||
The thing to note here, besides the transparency and the fact that it’s actually a good model that also works well on coding and agentic tasks, is that it’s the first release by a team formed less than a year ago, with a strong focus on iteration velocity. There’s more to come.

disclaimer: I‘m part of the training team, happy to answer any questions

ducktective 2 hours ago|
- Is it possible to train only on math and logic materials and expect the model's response in math questions to be superior to general models with the same training/inference compute hardware?

- Are there non-LLM approaches to the above task with the goal of achieving a non-hallucinatory agent?

tomComb 6 hours ago||
For a post to make such a big deal about sovereignty it is a bit misleading to not mention that the company is slated to be merged with Cohere, a Canadian company.

And that is a good thing - no need to hide it. Given the growing cost of keeping up, these few non-US, non-Chinese companies really need to do more sharing of efforts and costs.

Canada too is very much in need of sovereign AI options, but funding that on its own would be pretty much a waste of money. Would love to see this new German Canadian company cooperate with Mistral too, or maybe one of the Korean AI companies.

Loquebantur 5 hours ago||
Given that "sovereign" these days implicitly means independence from USA and China, and Canada being spiritually in the same boat as the EU, it's pretty spot on?

The costs of "keeping up" aren't really growing, on the contrary.

user_of_the_wek 3 hours ago||
I also read about Canada maybe joining the EU in some capacity, although I'm not sure if that was a joke. Eurovision could be a start.
mkesper 2 hours ago||
Ursula von der Leyen (president of european commission) propsed an "associated membership" with EU. Such an association is not defined anywhere in EU contracts yet. https://www.iss.europa.eu/publications/commentary/eus-first-...
ivo-42 4 hours ago|||
Hey there, I worked on Kolibri pre-training data. Agree with the point that non-US/non-Chinese labs need to pool effort. The merger with Cohere is public, nothing to hide and I'm also excited about it personally.
amoshebb 7 hours ago||
Qwen3.8 27B beats Kolibri 79.9 vs 70.8 in German in Kolibri's harness on Kolibri's benchmark.

Also, once the Cohere takeover is complete will they still be able to use this "sovereign" claim despite being 90% owned and 100% operated out of Toronto?

befelix 4 hours ago||
Disclaimer: I am part of the team that trained Kolibri, opinions are mine.

Qwen models have solid performance on benchmarks and we're transparent about this in the report. We're just happy to share a European alternative in the small model space, where I don't think we can afford to be fully dependent on China. We also put a lot of effort into German language quality things that don't show up in evals at all (style, Grammar, German reasoning, etc).

Sovereignty is imho mostly about choice and control over your data. Cohere is no different on that front and personally I'm quite excited about what we will build together post-merger.

slow_typist 46 minutes ago||
Souvereignty is about ownership and about knowing the training data. That is especially important given a business model aiming at the government as a key customer. With Schwarz, Cohere, SAP, NVIDIA and the like as backers society can’t have souvereignty in any meaningful way. Of course it would be a plus to keep US agencies out of the data streams. But this technology will be used to support and make decisions that impact citizens. That being said, incredible achievement.
andy99 6 hours ago|||
Yeah I think a model has to be actually good to claim sovereignty, as in competitive enough that people want to use it. Chinese and US LLMs are the only ones in these categories right now. Mistral and Cohere have the same problem, yes they are made in different countries but they are not competitive. They (France and Canada in this case) would be better off just downloading Chinese LLMs, even if they get cut off they still have the weights.

I’m most familiar with Canada, where sovereign is usually just an excuse to overpay someone connected for an inferior product with no strategic value.

Loquebantur 6 hours ago|||
Well, it's still independent from the USA and China.

The main problems with big corp AI are due to control of access in the first place and control of what they output.

When you make your industry reliant on such choke points, you render yourself the opposite of "sovereign" for sure. Having multiple independent suppliers at least mediates that.

bewareofscams 5 hours ago||
My account is banned, so for whomever with [showdead] on:

What's so "not-sovereign" for an open-weights American or an open-weight-open-training-process Chinese LLM?

Do we also need sovereign Linux (maybe), sovereign Postgres (most likely not), sovereign Python (def not)?

throw-qqqqq 12 minutes ago|||
You are not banned FYI
satvikpendem 1 hour ago|||
If it's banned then maybe you shouldn't be posting...

Given your other comments I can see why it was banned.

isusmelj 7 hours ago|||
I was also quite surprised to see that. Considering the effort Aleph Alpha put into to their new model, it seems like the Qwen team needs access to vast amounts of german data o.0 I'm really glad for these efforts for open models from within Europe.
BikDk 5 hours ago||
[dead]
niemandhier 8 hours ago||
I think at the moment the main thing a sovereign AI model needs to be good at is auditing the results of other models.

Right now one could run an open model for most government applications and it would be good enough, you just cannot trust any of these.

So having a sovereign controlled model audit the first one would basically act like a “trust adapter”.

If the second model is cheap and fast enough, there is a business model.

You don’t even need to audit all the intermediate steps, just tool calls and end results.

holaysuns 7 hours ago||
The issue with 'sovereign' is not about sneaky things a foreign entity might put in there, it's about control of the stack.

Since no individual buyer cares that much about geostrategic issues, it's almost impossible to get some random European company to think about buying anything other than 'whatever US or China' are making.

There has to be very concerted push to make a difference.

Legal mandates around sovereignty (be very careful here) could make a difference.

But there will also be market reaction: US/China companies will bend on some level and provide things like 100% EU hosted and even comply with some data source stuff.

The only real path is to be more competitive as a continent.

jwpapi 3 hours ago|||
If you self-host an open model on a sovereign hoster? They can only influence the responses correct? We are not talking about extracting information.

Or are we worried that open models send secret telemetry?

holaysuns 36 minutes ago||
We're talking about owning the stack - who gets to use what text, for what reasons, under what circumstances etc. Competitors, other nations, Russia / China, the 'global south', military contractors etc. etc..
Loquebantur 6 hours ago||||
The EU AI act isn't primarily about "geostrategy". It's about protecting society.

AI models act as a force multiplier for intelligence, in particular for generating information according to someone's wishes.

I.e. deepfakes and social media mis-/disinformation campaigns are a thing and having powerful AI allows you to do those at scales that can overwhelm society's resilience.

In general, even if you have "aligned" AI: aligned with whom or what?

Whom are you comfortable with lording as a some demi-god over you, dictating what to believe?

holaysuns 4 hours ago||
European regulations are 100% geostrategy and they are defending against external interests.

Those laws would not exist if European champions were leading the world.

"are you comfortable with lording as a some demi-god over you, dictating what to believe?"

Yes, Europe handed over all of the decisions about everything to foreign powers, now they have to enact regulations to try to constrain it.

Zuck et. al. make the investments decisions for Europe, by virtue of you all giving him the money and power to do that.

Stop giving him the money/power, then this regulation won't exist

Oras 7 hours ago|||
> Since no individual buyer cares that much about geostrategic issues

You clearly haven’t worked in enterprise in the EU. Location of data processor is the first thing they check. That’s why every major cloud has regions with different offerings, not just for HA and redundancy

holaysuns 7 hours ago||
They care because the are 'required to', otherwise those 'offerings' wouldn't even exist. It's mostly the result of regulatory action.
JaggerJo 6 hours ago|||
I’m working with clients (big companies) that care a lot - but not because they have to. They care because they lost trust in non EU players.
rokkamokka 6 hours ago||||
They also care because it's becoming increasingly risky to rely on US companies given the current political situation
bdangubic 6 hours ago|||
nope, that no longer holds any water. was true though couple of years ago…
holaysuns 4 hours ago||
The argument holds most of the water.

Yes - security concerns are now very real, but do not fall for this idea that anyone really cares about structural concerns.

Tons of EU companies are still selling crap to Russia, happy to look the other way wile the bear devours a neighbour, as long as profits are there.

What has happened is that Trump has given a face to the reality, and so companies are adjusting on some level.

But 1) it will be nominal 2) Trump will be gone and the impetus will fade 3) the conglomerates will react 4) modest legislation will mean ...

The can will get kicked down the road.

The invasion of E. Europe by Russia has not even caused defence spending or the size of Armies to change, doctrine is barely changing.

Nothing will make those beheamoths change other than other forms of structural concern.

The 'Riet of the Right' - partly due to collapse of Auto / China (among many other things) might cause some change. But the change won't happen until well after the damage has been done.

Brexit should have been an opportunity for institutional reform of the EU, but no - they blamed it all on populism.

Certain governing entities in Europe will try to move away from Microsoft, they will be pulled back.

There is hope, and certain champions can rise and causes pieces to collapse.

If SUSE had any true entrepreneurial whereiwthal - they would create a true consumer / prosumer / enterprise-user friendly variation and brand their flavour of Unix - and make sure that all Euopean governments us it exclusively, which would cascade into widespread use.

There are variations of insta, youtube, netflix that should all be Europe based that could 'theoretically happen' but it needs some structural impetus.

Things usually don't change. Usually there needs to be a collapse and re-order of the system for that to happne.

The people in charge just want to keep their jobs, their very high salaries, and protect their retirement, and will be happy to 'sell out' whatever other imperatives along the way.

Ending with a positive not - I would say current conditions mean the change is now 'plausible, it not likely' whereas before it was 'not very plausible'. So there's a candle, a bit of wind, but not a lot of tinder or dry wood.

mawadev 7 hours ago|||
That is a very good idea and there is a market for that, especially if a company manages to source the hardware and then install the stack on a german or EU customer site
adishiktensward 5 hours ago||
It depends on the use case, but I believe SLMs can do great with smaller tasks. People got so attached to 'general purpose' that they forgot software can be designed for specific, smaller use cases.
spijdar 8 hours ago||
The absence of any comparison to Qwen3.8 Flash, another MoE model with a small-ish (6B) number of active parameters, is pretty striking. Instead, it's compared with Qwen3-Next 80B-A3B, a model released almost a full year ago.

I get that doesn't invalidate the real "point" of the model, but...

bitexploder 5 hours ago||
That is what stood out to me as well. Qwen 3.8 Flash can run on very limited hardware as well and it is at least as good as Sonnet 5 in benchmarks like DeepSWE. With its n-gram design you can get flash next running on very limited GPU resources, as little as 16GB of VRAM.

People follow the latest frontier lab models with great attention and migrate to the next big model on their subscriptions. Meanwhile these local models have quietly gotten REALLY good. It is not even an exaggeration. It has happened in the last couple of months.

"Local model you can run at 40 t/s on a gaming machine that is better than Opus 4.6" is way less exciting than "OpenAI IS DOING CRIME!!! OpenAI SOLVED NAVIER STOKES. DARIO SAYS GLM 5.3 BAD! SLOW DOWN THE FRONTIER!".

(edit: also... totally ignore that 27B dense column over there where Qwen 3.8 27B beats Kolibri on nearly every single benchmark. Why would I choose to run this model?)

frumplestlatz 3 hours ago||
I think the simple reality is that if your goal is producing quality work output, you want the smartest possible model available.

What is the upper bound on the value of more intelligence applied to your problem domain?

bitexploder 2 hours ago||
The point of a model like this is aimed at providing local inference. You want the smartest model available, as long as it meets all of your other criteria. That isn't always maximizing on the absolute smartest model. There are many reasons to avoid a frontier model for now for a variety of reasons. All of this is even only relevant in the last 6 months anyhow. So it isn't like this is even some long term trade off or position I am proposing.
okamiueru 5 hours ago||
Can't you infer the comparisons you would like from baseline results provided? There are better results in the sibling post on HN: https://aleph-alpha.com/en/blog/kolibri-has-landed-a-soverei...
spijdar 4 hours ago||
Maybe?

Qwen3.8 27B scored notably higher in most of the provided benchmarks, including the German-specific ones. The only "downside" is that inference is much more costly and slow, since it's a dense model.

Qwen3.8 Flash-Next appears to usually "benchmark higher" than 27B, while remaining fast.

I'm sure I could dig up the equivalent benchmarks for Flash and do the comparison myself, but as far as inference goes, it's messy. Consider that Qwen3.5 35B-A3B scores higher than Qwen3.6 on some of the German-specific benchmarks.

So it seems superficially plausible that Qwen3.8 Flash-Next might not be "27B but faster" in the ways that are important for this model. Or it could just "be superior" in all ways.

Either way, I don't think an LLM has to be "the best" at anything to be worthwhile, necessarily. And I kind of distrust benchmarks on top of that, so...

martianvoid 9 hours ago||
I just tried to play around with it on my RTX pro 6000 setup, it spends way too many tokens on overthinking stuff even if it’s able to catch the correct approach

Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8

gizajob 9 hours ago||
> it spends way too many tokens on overthinking stuff

Yeah. It’s a German model.

sajithdilshan 8 hours ago|||
This joke is either gonna get over analyzed by the Germans or gonna fly past straight over their heads
hypfer 8 hours ago|||
I'm not sure if this is truly a good faith joke tbh.
jamiek88 3 hours ago||
[flagged]
hypfer 3 hours ago||
No, I am just standing my (and my nations) ground against what is simply bad faith disrespect and what could be described as casual racism (or at least spiritually equivalent to that, to not get into the racism definition debate).

You do not get to invalidate that.

People need to stop insulting my nation.

__

If the people you're joking with aren't laughing, you're not joking with them but about them, which is generally seen as a bad move, and something that we actually wanted to leave behind with all the social progress and all.

__

This really only happens because people (as all bullies) think that they would be free to do so, because the victims would not fight back.

(Un)fortunately though, the world now has bigger problems with genuine fascists, so neurotically cowering in fear and performatively saying "yes, punch me harder. Insult me more" can be stopped now.

__

I mean look at it. By now, it's at least better, but the first 30 comments or so were "haha germany bad haaa".

That's a terrible showing for the platform even if you're not german. What value was offered? Who would want this.

We don't even need to talk about nations and identities to see that that was just noise.

jijijijij 6 hours ago||||
Nothing flies anywhere, we use superior high speed trains. I faxed a Humorgenehmigungsantragantrag to our Bundesunterhaltungsministerium analyst, immediately. Laugh now, but you are merely lucky the reply got delayed. As soon as the leaves on the rails are dealt with, you will get a formal response that has washed itself! Let us see who is laughing then. It is nobody!
fph 3 hours ago||
But, most importantly, is Bundesunterhaltungsministerium a single token in Kolibri? This could make a huge difference for the model's performance in German.
jijijijij 1 hour ago||
Silly you, thinking the German economy is merit-based. The only question is, if the bribe fits a single token of appreciation, to win a tailored call for bids!
dudefeliciano 8 hours ago|||
petition to make this thread the official german jokes thread for this post, it's kind of annoying how everyone feels the need to make a top level comment for their oh-so-great "germans suck" joke.
tfburns 4 hours ago||||
Do models from other parts of the world underthink? :P
rubyfan 8 hours ago|||
Hey!
tzatzikyyy 8 hours ago||
Did you try lower reasoning modes as well?
dang 4 hours ago|
Related ongoing thread:

Aleph Alpha Kolibri: How the sovereign German LLM works - https://news.ycombinator.com/item?id=49943034 - Oct 2026 (162 comments)

More comments...