Aleph Alpha open-sources Kolibri-1: a 78.1B MoE with just 3.46B active params per token
lmoroney · x · 2026-10-06
Aleph Alpha released Kolibri-1 on Hugging Face under Apache 2.0, an English-German mixture-of-experts model with 78.1B total parameters and 3.46B active per token (about 4.4% of weights), up to 1M-token context (model card recommends serving at 262K or less), FP8 weights around 78 GB, pretrained on 768 B200 GPUs in Germany and Finland with over 20% German training data.
The author suggests using it as a hands-on MoE serving exercise: estimate memory from total params and per-token compute from active params, then pick hardware accordingly.
Related event: Aleph Alpha Open-Sources Kolibri: 78B-Parameter Model with 1M Context(3 posts)→
More from Models
- Gemini 4 Argon takes #1 on Text Arena leaderboard with Elo of 1,525 — jon_barron · 2026-10-06
- BaseDecision: a 0.4B local model wins 5 of 8 benchmarks on intent classification — OnlyFamousAddy · 2026-10-06
- Where is America's next great open-weights model? Reddit laments the HF leaderboard — Acrobatic-Laugh1856 · 2026-10-06
- Users still miss GPT-4o's intuition as newer models get wordy — flowersslop · 2026-10-06
- Sentdex: DeepSeek V4.1 Flash looks like a strong robotics control model, beating GLM 5.3 Flash — teortaxesTex · 2026-10-06
- Subscription "API value" math is inflated by token pricing, researcher points out — JeremyNguyenPhD · 2026-10-06