Aleph Alpha's Kolibri-1: 78B MoE with 3B active params, native FP8 weights
QuixiAI · x · 2026-10-05
Aleph Alpha released Kolibri-1 on Hugging Face, and a support request has been filed for the 39.5k-star open-source inference engine colibri. Key specs: FP8 (float8e4m3fn) weights in 128×128 blocks with dynamically quantized activations and FP8 KV cache; embeddings, LM head, norms and MoE router stay in bfloat16. It's an MoE with only 3B active of 78B total params, targeting SSD streaming on CPU/GPU/NPU with no weight conversion needed.
More from Infra
- Tiny decision model Jev classifies 1,000 papers for $0.08, letting frontier LLMs skip yes/no drudgery — TinfoilTricorn · 2026-10-06
- VC Funding Fell 66% in Last Hike Cycle, Putting ~25% of AI Lab ARR at Risk — menhguin · 2026-10-06
- TensorRT is 206% faster than llama.cpp in this local inference test — kalyan_kpl · 2026-10-06
- Consistent Hashing Explained: Why Modulo Assignment Reassigns Nearly Every Request — _jaydeepkarale · 2026-10-06
- CME launches compute futures as BlackRock's Larry Fink hails 'a new asset class' — Saul_Loveman · 2026-10-06
- Australia open-sources Matilda Jev, a 56.8ms decision model that skips text generation — Med1_Ai · 2026-10-06