Kimi K3 trains MoonViT-V2 from scratch to stabilize multimodal training
nrehiew_ · x · 2026-07-29
The thread highlights a key change in Kimi K3: the vision encoder, MoonViT-V2, is trained entirely from scratch with next-token prediction.
- Earlier practice, including Kimi K2.5, initialized the vision encoder from a contrastive model such as SigLIP.
- The authors say that once a pre-trained encoder is attached to the LLM, joint optimization becomes unstable.
- MoonViT-3D with SigLIP initialization shows persistent gradient spikes, while MoonViT-V2 stays more stable during training.
- The claim is that training with next-token prediction lets the encoder be shaped directly by the language-modeling objective.
- The post says MoonViT-V2 matches the SigLIP-initialized baseline on vision evaluations, suggesting contrastive pretraining may not be necessary at scale.
More from Models
- Macaron-V1-Tall trends on Hugging Face as a text-generation model — mindlab-research · 2026-07-29
- Claude Opus 5 is listed at $5 in and $25 out per million tokens — arena · 2026-07-29
- Kimi's Open Model Pricing Sparks Debate: The $20M Monetization Reality — BenBajarin · 2026-07-29
- User Finds Claude Opus Overly Verbose, Switches to Sonnet for Better Focus — brandon_galang · 2026-07-29
- Meta paper says RL can optimize code speed, with Qwen 2.5 7B and CWM 32B gains — burny_tech · 2026-07-29
- User reverses course and says GPT 5.6 Sol is actually a really good model — TheZachMueller · 2026-07-29