On October 3, 2026, German Reunification Day, Aleph Alpha released Kolibri, an open-weight language model for German and English under the Apache 2.0 license. It is a mixture-of-experts transformer with 78.1 billion total parameters and 3.46 billion active per token, a context window trained to 262,144 tokens and validated up to one million, and weights on Hugging Face. Kolibri is German for hummingbird, which suits a model whose pitch is being light on compute.
This post covers what Aleph Alpha says about Kolibri and how it holds up against its own tables, the details worth understanding (the tokenizer, German reasoning, trained abstention), the honest weaknesses, how to run it, and what the Hacker News discussion argued. It draws on Aleph Alpha's launch post and technical report, and on a detailed independent write-up by AI engineer Tejas Kumar, which includes his own tokenizer experiment.
TL;DR — what people are asking
| Question | Answer |
|---|---|
| What is it? | Aleph Alpha's German-English open-weight MoE model |
| Size? | 78.1B total, 3.46B active per token, 50 layers, 384 experts (6 active) |
| Context? | Trained to 262,144 tokens, validated to 1,048,576 |
| License? | Apache 2.0 for weights and configuration; Aleph Alpha keeps training code and methods |
| Training? | About 24T tokens, 768 NVIDIA B200 GPUs, Germany and Finland |
| German share? | 21.3% of pre-training tokens, about 4.3T |
| Strengths? | Math, long-context retrieval, German tokenizer, abstention, document QA |
| Weaknesses? | Closed-book knowledge, multi-turn tool calls, coding agents, mid-range long context |
| Beats a dense Qwen3.8 27B? | No: 70.8 vs 79.9 in German, 75.5 vs 80.2 in English (Aleph Alpha's table) |
| Memory? | About 78 GB of weights; all experts must be loaded |
The architecture in plain terms
Kolibri is built around efficiency choices that all trade parameters in memory for work per token.
- Mixture of experts. Each of 50 layers has 384 small experts plus one shared expert, and a router sends every token to 6 of the 384. That is how 78B parameters become about 3.5B of computation per token. Aleph Alpha says it chose many small experts over fewer wide ones because they tested better.
- Sliding-window attention. 40 of the 50 layers look only at the previous 512 tokens, and every fifth layer attends to the full context. This bounds decode cost and memory in most layers, which is what makes very long contexts affordable. Only the sliding-window layers encode position, which lets context stretch past the trained length without extra tricks.
- A size choice made on serving cost. The report compares 78B with a 123B variant: the larger model handled only 3 concurrent 256k-token queries on two H100s, while 78B handled 18 and decoded 28% faster.
- Training stability. Muon optimizer, no loss spikes reported, and a routing change called exact quantile balancing for expert load balance.
- Reasoning effort. A
reasoning_effortsetting of none, low, medium or high, so one model serves quick lookups and hard questions.
A comparison with the company's earlier model shows how fast the project moved:
| Kolibri Origin | Kolibri | |
|---|---|---|
| Finished pre-training | June 11, 2026 | September 11, 2026 |
| Total / active parameters | 30.6B / 3.27B | 78.1B / 3.46B |
| Pre-training tokens | 7.5T | 20T (about 24T with mid- and long-context stages) |
| Longest trained length | 65,536 | 262,144 |
| Experts (total / active) | 128 / 8 | 384 / 6 |
| Tokenizer vocabulary | 96,000 | 128,000 |
| Public release | None | October 3, 2026 |
Aleph Alpha says it hit 38 unplanned interruptions in 21 days of pre-training, about one per 10,000 GPU-hours, and that the pipeline recovered automatically from checkpoints at most 250 steps back. It also says it stopped and restarted Kolibri Origin's pre-training after finding a data-shuffling bug. That candor is part of why the technical report drew praise on Hacker News as unusually open.
The German-first details
A tokenizer built for compound words
German glues words into long compounds, and a tokenizer trained mostly on English slices them badly. Aleph Alpha trained a 128,000-token bilingual tokenizer with a new method it calls UniBPE: it keeps the bottom-up merging of byte-pair encoding but scores merges with the Unigram objective, which respects morphology. Its report measures German web text at 4.90 bytes per token for Kolibri against 4.35 for GPT-5's tokenizer, which works out to about 11.2% fewer tokens, with English compression roughly level.
Tejas Kumar ran his own test on the German constitution, the Basic Law, and its English translation:
| Tokenizer | German tokens | More than Kolibri | English tokens |
|---|---|---|---|
| Kolibri | 35,190 | 39,875 | |
| GPT-4o / GPT-5 tokenizer (o200k_base) | 41,482 | 17.9% | 39,737 |
| Qwen3.5 35B-A3B | 42,907 | 21.9% | 41,650 |
| Mistral Small 4 | 43,478 | 23.6% | 41,301 |
| Gemma 4 | 43,850 | 24.6% | 41,564 |
On that legal text Kolibri needed about 15% fewer tokens than the GPT-5 tokenizer, and tied it in English. One example from the report: the genitive "Bundessozialgerichtes" splits into four sensible pieces in Kolibri, but into six fragments in several competitors. Fewer tokens means cheaper and faster German, and more German fits in the context window.
Thinking in German
Reasoning models usually think in English even on German prompts. Aleph Alpha's earlier research found that a small amount of German reasoning data was worse than none, with a German math score dropping from 70.2 to 48.3 before recovering to 67.3 with much more data. Kolibri got the much more. It reasons in German on German prompts and posts the best German AIME 2025 score among models with about 3B active parameters, 87.5, against 84.4 for Nemotron 3 Nano.
Where the German data came from
German is 21.3% of pre-training, about 4.3T tokens over a 20T run. The unique German pool was 2.4T tokens: 1.3T from German web data the team curated with a pipeline retuned for German (standard filters removed the long-word prose that public administration writes in), about 1T rephrased from German documents by an LLM, and a small translated share. The report is explicit that heavy machine translation imports the cultural fingerprint of the source language: a corpus translated from English describes a German world that looks American.
Trained to say "I don't know"
Aleph Alpha trains abstention with a procedure it calls Merlin-Arthur, a three-player game. Arthur is the model. Merlin supplies a context that makes the right answer easier to find. Morgana strips out the evidence the answer depends on and tries to lure Arthur into a guess. Arthur cannot tell which one he faces, so the only winning strategy is to check whether the evidence supports an answer. Wrong guesses, even lucky ones, are penalized when evidence is missing.
The reported results: Kolibri abstains instead of answering wrong on 44% of AA-Omniscience items, against 15% for Kolibri Origin. In Kumar's reading, comparators were 11.1% for Qwen3.5 35B-A3B and 23.7% for GPT-OSS 120B, with only Qwen3.6 35B-A3B higher at 56.7%. Aleph Alpha's AA-Omniscience Index puts Kolibri at −32.8, ahead of Qwen3.5 35B-A3B at −47.3 and GPT-OSS 120B at −35.2 but behind Qwen3.6 35B-A3B at −15.3.
The numbers, and how to read them
These come from Aleph Alpha's table, run through its own harnesses at the highest reasoning effort for each model.
| Benchmark | Kolibri | Qwen3.5 35B-A3B | Nemotron 3 Nano | Qwen3.8 27B (dense) |
|---|---|---|---|---|
| Overall, English | 75.5 | 74.7 | 65.6 | 80.2 |
| Overall, German | 70.8 | 69.8 | 59.3 | 79.9 |
| AIME 2025, English | 96.9 | 88.1 | 89.6 | 97.9 |
| AIME 2025, German | 87.5 | 76.7 | 84.4 | 96.5 |
| GPQA Diamond, English | 84.3 | 83.8 | 73.9 | 89.2 |
| LiveCodeBench v6 | 85.9 | 77.8 | 71.3 | 93.8 |
| SWE-bench Verified | 66.4 | 71.6 | 38.6 | 72.6 |
| Terminal-Bench 2.1 | 27.7 | 39.7 | 9.7 | 76.8 |
| BFCL v3 multi-turn | 39.8 | 54.0 | 47.9 | 42.5 |
| AA-Omniscience accuracy | 14.8 | 22.0 | 19.5 | 17.5 |
Three honest readings follow.
- Among MoE models with about 3B active parameters, Kolibri leads overall. The margin over Qwen3.5 35B-A3B is about one point in each language.
- A dense 27B model is simply better. Qwen3.8 27B beats Kolibri by about 5 points in English and 9 in German, but it activates roughly eight times as many parameters per token. That is the trade Kolibri is built around: cheaper to serve per token, but heavier to hold in memory.
- It is not a coding or tool-use model. Terminal-Bench, SWE-bench and multi-turn function calling all trail the field.
Aleph Alpha also reports internal "customer-proxy" suites across five sectors, such as automotive supply, semiconductors, public sector, industrial drives and aerospace, with scores climbing across training checkpoints (for example aerospace from 0.14 to 0.59 between the earlier model and Kolibri). These are the company's own benchmarks on documents the model did not see in training, and they cannot be independently checked.
Some Hacker News readers noted that the comparison set does not include the very newest competitors, such as Qwen3.8 Flash, and that Qwen3.8 27B beats Kolibri in the same table. Both are fair points about framing: the table is honest, but "best in its class" depends on where you draw the class.
Weaknesses
Aleph Alpha publishes its weak rows alongside the strong ones. In summary:
- It knows less from memory. Last of the 12 compared models on closed-book questions, and 14.8% accuracy on AA-Omniscience against 22.0% for Qwen3.5 35B-A3B. It knows when it does not know, but it also knows less.
- Multi-turn tool calling is weaker than several peers.
- Coding agents lag.
- Long context is mixed. The base model leads at 1M tokens on RULER (63.2 vs 57.5 for Qwen3.5 35B-A3B's base) but trails at 128K tokens (67.9 vs 89.9).
- Memory. All 78B parameters must be loaded.
- Two languages. A deliberate "depth over breadth" choice, which some readers found limiting since multilingual models transfer knowledge across languages.
- Overthinking. One Hacker News user running it on a single workstation GPU reported it spending many tokens deliberating, though decoding was fast at around 170 tokens per second at 8-bit. A single user report, not a measurement.
How to run it
You need the aleph-alpha-inference package, which provides the vLLM plugin and installs the vLLM version it supports, or the project's container image:
pip install "aleph-alpha-inference>=1.0"
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
For contexts beyond 262,144 tokens, add --max-model-len 1048576 and --hf-overrides '{"max_position_embeddings": 1048576}'. The recommended sampling settings are temperature 1.0, top_p 0.97 and top_k 128. The server is OpenAI-compatible, and reasoning effort is passed through the chat template arguments.
Hardware: roughly 78 GB for the weights at 8-bit, for example two 80 GB GPUs, or one H200, B200 or B300. That has a practical consequence for local-AI buyers: a 128 GB unified-memory machine can hold the weights with limited room for context, while a 64 GB machine cannot; see our coverage of the new 64 GB DGX Spark. One reader reported that the Hugging Face repository initially offered only a higher-precision format that would not fit on a Mac, and community quantizations had not appeared yet, so check the model card for formats and expect GGUF builds later.
If you want to run open models locally with an agent, our guides to Qwen on llama.cpp with OpenCode and choosing open-weight versus closed models cover the surrounding decisions.
When Kolibri is the right pick
- Regulated German-language document work on your own hardware: public administration, finance, aerospace, manufacturing. Retrieval over long German laws, contracts and manuals plays to the tokenizer, the long context and the abstention training.
- Workflows where a wrong confident answer is worse than no answer. Pair it with retrieval; see our guide to grounding, RAG and fine-tuning.
- Teams that need an Apache 2.0 model with a documented training pipeline.
It is the wrong pick for coding agents, closed-book trivia, languages beyond German and English, or anyone who cannot spare the memory. In those cases a dense Qwen model or a smaller specialist wins.
The sovereignty argument
Aleph Alpha defines sovereignty two ways: built in Germany and trained on infrastructure in Germany and Finland, under European and German law with no foreign control; and delivered to customers as open weights with freedom of deployment and IP safety. It has signed the EU's General-Purpose AI Code of Practice and says it built Kolibri with the EU AI Act and GDPR in mind.
The Hacker News thread tested those claims:
- Competitiveness. Several commenters argued sovereignty is only credible if the model is good enough to use, and that a model worse than Qwen3.8 27B is a hard sell. Others countered that sovereignty is a goal in itself, with quality following, and that having an independent supply matters even when it is behind.
- Ownership. Commenters noted Aleph Alpha is merging with Cohere, a Canadian company, and one said the combined company would be mostly owned and operated from Toronto. We cannot verify ownership percentages; our earlier report describes the merger without them. See Cohere and Aleph Alpha merge.
- Training data. One commenter asked how "intellectual-property safety" squares with training on web data; German web text was curated from Common Crawl, so the claim is about process and transparency, not an absence of web scraping. No independent audit exists yet.
- Foreign tools in the pipeline. As relayed from the model card, rephrasing used models from outside Europe, such as Gemma 4 and Mistral-NeMo, and quality filters used Qwen-labeled data. That is common practice, but it complicates a purist definition of sovereign.
- Openness. Several readers praised the technical report's detail about data and process as unusually open, though others noted the data itself is not released.
For a structured way to think about all this, read our explainer on what sovereign AI actually means.
What people are asking
Is Kolibri the best German LLM?
Among open MoE models with a few billion active parameters, it has the top German score in Aleph Alpha's evaluation. A dense Qwen3.8 27B scores higher. "Best" depends on whether you optimize for serving cost per token or total quality.
Is Kolibri open source?
The weights and configuration files are Apache 2.0. Aleph Alpha keeps rights to its training code and methods, and the training data is not released, so it is open-weight rather than fully open.
Can I fine-tune it?
The license allows it. The practical barrier is memory: fine-tuning a 78B-parameter MoE needs substantial hardware, and parameter-efficient methods will likely arrive with community tooling.
Why only German and English?
Aleph Alpha calls it a deliberate choice of depth over breadth. Whether that beats a multilingual model for your use case is an empirical question to test on your documents.
Honest limitations
- Benchmark numbers are from Aleph Alpha's own evaluation and technical report; we have not reproduced them or run the model.
- The tokenizer test is Tejas Kumar's, run on one document; we did not repeat it.
- Details attributed to the model card, such as training-data tooling, are as relayed by that write-up.
- Hacker News comments are individual opinions, and the ownership claim is unverified.
- Hardware figures come from the model card as relayed; confirm formats and memory needs in the repository.
Bottom line
Kolibri is a credible, openly licensed, German-first model with real engineering behind it: a German-aware tokenizer, German reasoning, trained abstention, and a documented pipeline. It leads its efficiency class, loses to a larger dense model, and is not a coding model. If your problem is long German documents on your own servers, it deserves a place in your evaluation set. If your problem is anything else, test before you commit.
Related on explainx.ai
- What is sovereign AI? Layers, models and trade-offs
- Cohere and Aleph Alpha merge
- Apertus: a fully open sovereign foundation model
- Europe AI landscape: sovereign compute and the EU AI Act
- Choose open-weight vs closed AI models
- Grounding, RAG or fine-tuning
- Qwen 3.8 27B open-weight comparison
- Kimi K3 architecture notes
Sources: Aleph Alpha, "Kolibri Has Landed: A Sovereign Open-Weight Model" (October 3, 2026) and technical report · Tejas Kumar, "Aleph Alpha Kolibri: How the Sovereign German LLM Works" · Hacker News discussion, October 3, 2026
Figures reflect Aleph Alpha's published materials and secondary reviews as of October 3, 2026.
