DEV Community

Cover image for Aleph Alpha Kolibri 1: open German-English MoE model
techaiwire
techaiwire

Posted on Originally published at techaiwire.com

Aleph Alpha Kolibri 1: open German-English MoE model

German AI company Aleph Alpha released Kolibri 1 on October 3, 2026, an open-weight language model built for German and English. It has 78.1 billion parameters, but only 3.46 billion do work on each token, which keeps it fast for its size. The weights are free to download and use commercially under the Apache 2.0 license. Aleph Alpha pitches it at governments and regulated companies that want to run AI on their own hardware instead of sending documents to an outside provider.

What Kolibri 1 is

Kolibri is a mixture-of-experts (MoE) model. Instead of one large network, it has many smaller "experts," and a router picks a few of them for each token. According to the model card on Hugging Face, it has 50 layers with 384 experts each. Six routed experts plus one shared expert handle every token.

Spec Kolibri 1
Total parameters 78.1 billion
Active per token 3.46 billion
Layers / experts 50 layers, 384 experts per layer
Context window Up to 1,048,576 tokens
Vocabulary 128,000 tokens
Weights FP8, about 78 GB
License Apache 2.0
Knowledge cutoff June 18, 2026

The context figure needs a note. The model card lists a native context of 1,048,576 tokens, about a million. Its longest training stage used 262,144 tokens, and RuntimeWire and the tej.as blog describe 262,144 as native, with the full million tested.

How it was trained

Aleph Alpha trained Kolibri on 768 Nvidia B200 GPUs for 21 days, using data centers in Germany and Finland, its announcement says. The model card counts about 23.6 trillion training tokens in total, starting with 20 trillion in pre-training.

The sources differ on how much of that is German. Aleph Alpha's blog says 21.3%, about 4.3 trillion tokens. The model card gives 23.9% of the pre-training data. Either way, more than a fifth of the training data was German.

That focus shows up in the tokenizer, the part that splits text into pieces. RuntimeWire reports that Kolibri breaks Germany's Basic Law into 35,190 tokens. The tokenizer used by GPT-4o and GPT-5 needs 41,482 for the same text. Fewer tokens means lower cost and more room in the context window.

How it scores

Aleph Alpha and the tej.as blog list these results:

Benchmark Kolibri 1 Comparison
AIME 2025 (English) 96.9%
AIME 2025 (German) 87.5% Nemotron 3 Nano: 84.4%
Overall (German) 70.8% Qwen3.5 35B-A3B: 69.8%
Overall (English) 75.5%
GPQA Diamond (English) 84.3%
HumanEval+ 92.7%
Long context, 1M tokens 63.2% Nemotron 3 Nano: 57.5%

Aleph Alpha also says it trained the model to admit gaps. "Kolibri is trained to say 'I don't know' when the answer isn't in the context," the company writes. According to tej.as, the model also "reasons in German on German prompts."

Who it is for

Aleph Alpha aims Kolibri at "sovereign, mission-critical" work in regulated fields, such as public administration, aerospace and manufacturing. The pitch is control: a customer can run the model on its own servers under German and European law. The tej.as post sums up the argument as "full freedom of deployment and intellectual-property safety."

Kolibri arrives in a busy week for open models. Ai2 released AstaBrief 8B, a small model for cited research reports, on October 2.

What this means for developers

Plan your hardware around the full 78 GB, not the 3.46 billion active parameters. As tej.as puts it, "only about 3.5 billion parameters work on each token, but all 78 billion have to sit in memory." The model runs fast once loaded, but you need more than 78 GB of accelerator memory for the FP8 weights alone, plus room for the context.

If your product serves German speakers, test Kolibri against your current model on real German text. The tokenizer savings alone can cut costs, and its German scores edge out the two open models tej.as compares it with. Measure answer quality on your own documents, not just benchmarks.

The Apache 2.0 license makes it easy to adopt. You can fine-tune it, ship it inside a product, and run it fully offline. For teams in the public sector or in healthcare, that removes the data-transfer question that blocks many hosted models.

Treat the million-token context with care. The longest training stage used 262,144 tokens, and the 1M-token score of 63.2% shows recall drops at full length. Keep critical material within a quarter-million tokens until you have tested longer inputs on your own workload.


This article was first published on Tech AI Wire.

Also available in

Deutsch · 日本語 · Français · Español · Português

Related on Tech AI Wire

Sources

Top comments (0)