deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Aleph Alpha releases Kolibri, an open-weight 78B-parameter MoE model for German and English

German AI firm Aleph Alpha has released Kolibri, an Apache 2.0 mixture-of-experts model with 78 billion total parameters but only 3.46 billion active per token, trained on infrastructure in Germany and Finland.

Aleph Alpha releases Kolibri, an open-weight 78B-parameter MoE model for German and English

Aleph Alpha ships Kolibri under Apache 2.0

On 3 October 2026, German AI company Aleph Alpha released Kolibri, an open-weight language model for German and English whose weights are available on Hugging Face under an Apache 2.0 license. According to the company's launch materials, summarised in a technical write-up on the tej.as blog that reached Hacker News's front page, the model was trained from scratch on infrastructure in Germany and Finland, a point Aleph Alpha leans on heavily in its pitch for sovereign European AI. The name is German for hummingbird, a fitting label for a model whose trick is staying light.

The headline numbers: 78.1 billion total parameters arranged as a mixture of experts, with only 3.46 billion, about 4.4%, active for any given token; a native context window of 262,144 tokens, validated up to 1,048,576; tool calling; four reasoning levels; and a knowledge cutoff of 18 June 2026. Training ran over roughly 24 trillion tokens, more than a fifth of them German, on 768 NVIDIA B200 GPUs.

Built around German text

Kolibri's 50 layers each contain 384 experts plus one shared expert, and a router sends every token to 6 of the 384. That is what turns 78 billion parameters into about 3.5 billion of actual computation per token. The trade-off, which the model card states plainly, is that the whole parameter set must sit in memory even though only a fraction fires at any moment; in FP8 the weights take up about 78 GB.

The tokenizer is the more unusual part. Aleph Alpha trained a 128,000-token vocabulary with a new algorithm it calls UniBPE, which keeps byte-pair merging but scores candidate merges with the Unigram objective so that German compound words survive intact. The company's technical report claims 11.2% fewer tokens for German text than GPT-5's tokenizer, the best result among nine alternatives it measured.

The tej.as author tested this independently by tokenising the entire German constitution, 185 KB of dense legal text, with six tokenizers. Kolibri needed 35,190 tokens against 41,482 for o200k_base, the tokenizer behind GPT-4o and GPT-5, roughly 15% fewer, while in English the two were effectively tied at 39,875 versus 39,737. Fewer tokens mean cheaper inference on German input and more German text fitting in the same context window.

Attention is split as well: 40 of the 50 layers use a 512-token sliding window, while every fifth layer attends to the full sequence. Position information exists only in the sliding-window layers, which is what lets context stretch past the trained length without extra tricks. At one million tokens on the RULER long-context benchmark, Kolibri's base model scored 63.2 versus 57.5 for the base model of Qwen3.5 35B-A3B.

Sovereign, with an asterisk

Aleph Alpha frames sovereignty two ways: the model was built by teams in Germany, trained under European and German law on hardware in Germany and Finland, with, in the company's words, "no foreign control"; and customers get freedom over deployment and intellectual-property protection, so a ministry or a car supplier can run it on its own servers with data never leaving the building. Aleph Alpha has also signed the EU's General-Purpose AI Code of Practice, says the model was designed with the EU AI Act in mind from the start, and claims its own evaluations place Kolibri above every compared model of its size in both languages.

The model card is candid, though, that non-European models touched the data pipeline: English web text was rephrased with Google's Gemma 4, German text with Mistral-NeMo, and Qwen3-32B labelled data for the quality filters. Aleph Alpha says it then filtered the training data for the political bias such models can carry. The license has a boundary too: Apache 2.0 covers the weights and configuration files, while the training code and methods remain with the company.

Why it matters

European AI sovereignty has mostly been argued about in policy documents; Kolibri is a concrete artifact of it, with open weights, a permissive license, deliberate German-first engineering, and documentation aligned with EU rules. For German-language workloads, the tokenizer results and the long-context benchmark point to genuine efficiency gains rather than marketing, and the roughly 78 GB memory footprint is manageable for exactly the on-premises deployments that want data sovereignty. At the same time, the training pipeline shows what sovereign means here: control over law, infrastructure and deployment, not isolation from foreign models. That distinction matters for anyone assessing Kolibri as a fully European stack.

  • #aleph-alpha
  • #kolibri
  • #mixture-of-experts
  • #open-source
  • #sovereign-ai

Related posts