Kimi K3’s 1.6 TB weights may hide 20–40T training tokens
johnseach · x · 2026-07-28
The post frames Kimi K3 as a massive open model: 2.8 trillion total parameters, 104B active parameters, and native MXFP4 weights that weigh about 1.56–1.6 TB. It then asks how much data the model likely trained on, since Moonshot has not disclosed the training set size.
Based on Kimi K2’s reported 15.5T tokens, K3’s larger active parameter count, and claimed scaling efficiency gains, the author estimates a realistic training range of 20–40T tokens. Converting that to storage, they argue the curated training corpus could be roughly 80–160 TB, meaning the weights are only a compressed snapshot of a much larger training pipeline.
More from Models
- Tiron ships as an open-weights model for multi-speaker meeting transcription — Balance- · 2026-07-28
- Dev Team Drops Anthropic Max for Codex and China’s Kimi, GLM, Grok in Cursor — haider1 · 2026-07-28
- Anthropic’s Claude Opus 5 gets an official prompting guide buried in the API docs — JarnoDuursma · 2026-07-28
- DeepSeek V4 GA rumors point to NDA-heavy rollout and weeks of black-box release — teortaxesTex · 2026-07-28
- Reddit user asks whether KIMI-K3 stays uncensored through OpenRouter — Suhan_XD · 2026-07-28
- Kimi K3’s 2.8T open model puts pressure on Anthropic’s $200 plan — haider1 · 2026-07-28