2026-10-03

Aleph Alpha Releases Kolibri, a 78B-Parameter Mixture-of-Experts Model With 3.46B Active Parameters, as Apache 2.0 Open Weights

AIModelsOpen Source🌍 Europe

Aleph Alpha released Kolibri on October 3, 2026, an English-German Mixture-of-Experts model with open weights under Apache 2.0. Aleph Alpha's announcement put it in one line: "Small bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now the weights are yours." About 4.4% of the model is active per token. It was built end to end in Germany, trained in Germany and Finland, and documented in a 189-page technical report.

Two checkpoints are available, Kolibri-1 in FP8 (about 78 GB) and Kolibri-1-BF16 (about 156 GB). Context is 1,048,576 tokens, with the card recommending up to 262,144 for serving efficiency, plus tool calling and an explicit reasoning mode.

How it was built

Pre-training covered about 24T tokens in three stages: 20T at 16k context, 3.4T of mid-training at 64k, and 200B for extension to 256k. Trending Topics reports the 20T-token stage ran on 768 Nvidia B200 GPUs for about 21 days.

German is over 20% of the data, including more than 2T German tokens curated from the web or generated synthetically. Its 128k-token tokenizer needs 11.2% fewer tokens for German than GPT-5.

The architecture has 50 blocks, 40 with sliding-window attention and 10 with full attention. Each Mixture-of-Experts layer routes a token to 6 of 384 experts, plus one shared expert. Post-training used supervised fine-tuning on 537B tokens, including German reasoning, then reinforcement learning on more than 1.2M internally curated tasks covering reasoning, tool use, code and retrieval.

Experiments ran through Model Factory, a version-controlled "training as code" system (Savanna). The authors say decision speed, not compute, was the bottleneck; Kolibri Origin became Kolibri in about three months.

Results and weak spots

The report places Kolibri on the quality-versus-serving-cost Pareto frontier in both languages, claiming 1.6 to 2.7x throughput and 20-plus points of quality over some compared open models. One chart shows about 47,000 bytes per second at roughly 71% quality, against about 34,500 (70%) for GPT-OSS and about 38,500 (67%) for Qwen3.6 A3B. The comparison set includes Qwen3.6-35B-A3B, Nvidia Nemotron 3 Super 120B-A12B and Mistral Small 4 119B-A6B.

Kolibri beats every compared Mixture-of-Experts model of its size on the English and German aggregates. Maths is the standout: it reportedly scores 96.9 on AIME 2025 and 96.0 on AIME 2026, and the report says it matches or beats models with 3x more active parameters on maths and code. The dense Qwen3.8 27B scores higher overall but activates about 8x more parameters. Hallucination drops markedly versus Kolibri Origin.

Closed-book factual recall is weak: Kolibri ranks last of twelve on RGB Closed-Book and also trails on AA-Omniscience and RGB Fact-Check. The multi-turn BFCL tool-calling splits are another soft spot. Trending Topics, under the headline "Aleph Alpha's Kolibri Is No Match for the Open-Weight Leaders", notes that the comparison models are about six months old and today's open-weight leaders are far ahead.

Sovereignty and compliance

Kolibri targets regulated, sovereign deployments such as public administration, industry and aerospace. At 3.46B active parameters it is far cheaper per token to serve than the memory-hungry releases in Open Weights You Can't Actually Run, and Apache 2.0 is a permissive rung of the ladder in How Open Is 'Open'?. The Hugging Face summer report put the open-weights ceiling in Chinese hands in nearly every month of 2026.

The pipeline is designed around the EU AI Act, the GPAI Code of Practice and GDPR (see the Digital Omnibus changes). It filters pirated and harmful content, redacts personal data, and trains the model to abstain when the context does not support an answer. Pages 111 to 129 are a legal deep dive on copyright, covering the EU text-and-data-mining exception and the AI Act. Next steps are better German reasoning, fewer hallucinations and more European languages.

Kolibri follows Soofi S as a second German sovereign open model. Jonas Andrulis stepped down as CEO and became Chairman in October 2025. TechCrunch reported on April 25, 2026 that Cohere would merge with Aleph Alpha at roughly a $20 billion combined valuation, and a definitive agreement was signed on September 16, 2026. Pending regulatory approval, closing is expected later in 2026 under the Cohere brand, with Toronto and Berlin as dual headquarters and Heidelberg as a research centre. France's economy minister has said European AI cannot rest on Mistral alone, and Mistral is selling regional inference and trust.

Read next