Posted by bastitx 16 hours ago
additional paper: https://tej.as/blog/aleph-alpha-kolibri
Of course, it's impossible to know for sure what was LLM processed or not, but this post did get classified that way. That's why the software flagged it.
That explains why they're framing it around sovereign deployment for public administration / aerospace rather than chasing general English benchmarks. The Pareto-frontier claim on throughput vs quality (Figure 1, 8xB200 evals) is also refreshingly honest — serving cost matters a lot for regulated on-prem use.
Would love to see more detail on how the synthetic German data was validated for quality, and how MergeMix data mixing affected German vs English trade-offs. Apache 2.0 open weights is a big plus here.