Kimi's minor bump was a full pretrain: two-phase training splits omni data from the pro model
stochasticchasm · x · 2026-09-22
Analyzing Kimi's latest version, the author flags two surprises:
- A minor version bump involved a full new pretrain with the same architecture and shapes — arguably more of a "data version bump."
- Training switched from single-phase omni-modality (as in k3) to two-phase, and the big model barely received the phase-2 omnimodal mix, a notable data mismatch between sizes.
The author infers product positioning: the flash tier is built as a Gemini-Flash-like multimodal data processor, while the pro tier is a SWE-bench-maxxing coding model.
Related event: Analysts spot mid-training optimizer switches from Adam to Muon(3 posts)→
More from Models
- DFlash2 reportedly chokes on super long prompts when paired with GLM 5.3 Flash — TheZachMueller · 2026-09-22
- Shots fired at DeepSeek: MiMo pioneered MOPD, observers weigh in on V2.6 — teortaxesTex · 2026-09-22
- Jev fails counting r's in strawberry, aces it 168/168 when given letters — BLUECOW009 · 2026-09-22
- Blogger corrects himself: the real surprise is MiMo-2.6, beating Grok 4.7 at much lower cost — kimmonismus · 2026-09-22
- Jev: InstructGPT coauthor Diogo Almeida on System One models for prod, not God — Latent Space · 2026-09-22
- Unverified DataBench Charts Fuel Rumors of OpenAI's Internal Model 'Luna' Ahead of GPT-6 — almmaasoglu · 2026-09-22