Peter Gostev debunks model sparsity leak: Kimi 26:1, DeepSeek 32:1, 1.2T active params implausible
inductionheads · x · 2026-09-07
Peter Gostev pushed back on a viral leak claiming an OpenAI model with 1.2T active parameters, running the MoE sparsity math: Kimi k3 is 2.8T total with 104B active (26:1), and DeepSeek v4 is sparser still at 1.4T total to 49B active (32:1). Given the industry trend toward even higher sparsity, he argues there's no plausible way OpenAI would sit at 5:1 or 10:1 — making the 1.2T claim clearly fabricated.
More from Models
- Next-gen model training could hit 5x pretraining compute if 300k GB200 rumor holds — scaling01 · 2026-09-07
- GPT-6 could use 2-5x pre-training compute for RL, speculates scaling01 — scaling01 · 2026-09-07
- Banned triton is reward hacking; unbanned PyPI shortcut is a task bug — xeophon · 2026-09-07
- Frontier pre-training runs likely capped at ~2 months to avoid wasting algorithmic progress — scaling01 · 2026-09-07
- Codex recreates a full game in 40 minutes from one vague prompt, no Blender skills needed — FuSheng_0306 · 2026-09-07
- Vincent Conitzer shows what GPT-6 Astra can't do, pushing back on loose AGI labels — conitzer · 2026-09-07