Kolibri by the Numbers: 492,000 GPU-Hours, ~$3M Training Run — 'Large Enough' Beats Frontier
julsimon · x · 2026-10-06
Julien Simon's Substack deep-dive on Aleph Alpha's Kolibri (78B params, German/English, Apache 2.0, trained from scratch): 392,000 GPU-hours pre-training on 768 NVIDIA B200s over 21 days, 492,000 total GPU-hours, $3M training cost with American chips and Chinese teacher models. The paradox: a company that said a European LLM alone isn't a business, now merged-into-Cohere, shipped one anyway — because from-scratch 'good enough' has collapsed in price. Not a DeepSeek moment; the point is eligibility, a competent model European ministries can buy.
Related event: Ex-AWS Evangelist Debunks Kolibri as Europe's DeepSeek Moment(2 posts)→
More from Infra
- Hugging Face Kernels quickstart: load GPU-optimized kernels in one line — ariG23498 · 2026-10-06
- ZML inference already runs on Tenstorrent, Qualcomm, Intel and FuriosaAI chips — RemiCadene · 2026-10-06
- NInfer6000 hits 400 tok/s decode on RTX6000 running Qwen 3.8 Flash Next — lkarlslund · 2026-10-06
- Drax datacentre would burn 4.9m tonnes of wood a year, emissions near double Gatwick flights — nordicinst · 2026-10-06
- Singapore data center operator DayOne files for US IPO after H1 revenue tripled to $512M — zephyr_z9 · 2026-10-06
- Strata claims 6GB VRAM can match RTX 5090-level inference, with full Qwen4 support planned — lxfater · 2026-10-06