Aleph Alpha's Kolibri-1: 78B MoE with 3B active params, native FP8 weights

QuixiAI · x · 2026-10-05

Aleph Alpha released Kolibri-1 on Hugging Face, and a support request has been filed for the 39.5k-star open-source inference engine colibri. Key specs: FP8 (float8e4m3fn) weights in 128×128 blocks with dynamically quantized activations and FP8 KV cache; embeddings, LM head, norms and MoE router stay in bfloat16. It's an MoE with only 3B active of 78B total params, targeting SSD streaming on CPU/GPU/NPU with no weight conversion needed.

Related event: Aleph Alpha open-sources Kolibri-1: a 78B MoE with only 3.46B active params(14 posts)→

Original post →

More from Infra

Infra channel →