Aleph Alpha Releases Kolibri: a From-Scratch 78B MoE With Only 3.46B Active Params, Apache 2.0
solyarisoftware · x · 2026-10-04
Aleph Alpha released Kolibri, a MoE model trained from scratch (not a Qwen fine-tune) under Apache 2.0. Key specs: 78.1B total parameters with only 3.46B active per token (4.4%), 384 routed experts (6 selected + 1 shared), 256K native context validated to 1M, German/English bilingual, tool calling, and none-to-high reasoning levels. On the company's own benchmark panel it outperforms Qwen3.6-35B-A3B on AIME 2026 (96.0 vs 91.0), GPQA Diamond (84.3 vs 83.4), and LiveCodeBench v6 (85.9 vs 82.5). Caveat: the official FP8 weights weigh 78GB, so serving still requires substantial memory.
More from Models
- Daniel Han publishes summary of LLM benchmarks you can actually trust — danielhanchen · 2026-10-06
- Claim Verification Benchmarks Mostly Test Retrieval, Not Reasoning, Finds 24K-Trace Study — deliprao · 2026-10-06
- COLM26 study: LLMs ace claim verification benchmarks by taking shortcuts, not verifying — deliprao · 2026-10-06
- Opus 5.5 uses 26k tokens vs Astra's 12k yet costs 23% less per task at equal AA score — ChrisGPT · 2026-10-06
- GPT-6 Astra claimed to be first AI crossing world-class astrophysics threshold — johnseach · 2026-10-06
- $500/mo AI subscription is huge money in Jakarta: PPP pricing debate — sujingshen · 2026-10-06