Aleph Alpha Releases Kolibri: a From-Scratch 78B MoE With Only 3.46B Active Params, Apache 2.0

solyarisoftware · x · 2026-10-04

Aleph Alpha released Kolibri, a MoE model trained from scratch (not a Qwen fine-tune) under Apache 2.0. Key specs: 78.1B total parameters with only 3.46B active per token (4.4%), 384 routed experts (6 selected + 1 shared), 256K native context validated to 1M, German/English bilingual, tool calling, and none-to-high reasoning levels. On the company's own benchmark panel it outperforms Qwen3.6-35B-A3B on AIME 2026 (96.0 vs 91.0), GPQA Diamond (84.3 vs 83.4), and LiveCodeBench v6 (85.9 vs 82.5). Caveat: the official FP8 weights weigh 78GB, so serving still requires substantial memory.

Related event: Aleph Alpha open-sources Kolibri-1: a 78B MoE with only 3.46B active params(14 posts)→

Original post →

More from Models

Models channel →