Kimi's minor bump was a full pretrain: two-phase training splits omni data from the pro model

stochasticchasm · x · 2026-09-22

Analyzing Kimi's latest version, the author flags two surprises:

The author infers product positioning: the flash tier is built as a Gemini-Flash-like multimodal data processor, while the pro tier is a SWE-bench-maxxing coding model.

Related event: Analysts spot mid-training optimizer switches from Adam to Muon(3 posts)→

Original post →

More from Models

Models channel →