Aleph Alpha has released Kolibri, a large open-weight language model designed for English and German. The model uses a Mixture-of-Experts architecture with 78.1B total parameters, yet activates only 3.46B parameters for each token. This sparse activation is intended to keep inference costs and latency closer to those of a much smaller model while retaining the capacity of a large one.

Kolibri also features a 1M-token context window, allowing it to process very long documents or conversations in a single pass. The model supports per-request reasoning effort, letting users trade compute for more thorough responses depending on the task. Weights are released under the Apache 2.0 license in FP8 precision, and the model is designed to run on a single Nvidia B200 or H200 GPU.

The release stands out for combining bilingual capability, long context, and permissive licensing in a package that fits on one high-end accelerator. By keeping active parameters low, Aleph Alpha positions Kolibri as a practical option for organisations that need German-language understanding without the infrastructure overhead of a fully dense 78B model.