Kimi K3 trains MoonViT-V2 from scratch to stabilize multimodal training

nrehiew_ · x · 2026-07-29

The thread highlights a key change in Kimi K3: the vision encoder, MoonViT-V2, is trained entirely from scratch with next-token prediction.

Related event: Kimi K3 case studies show kernel optimizations, a Triton-like compiler, and a chip prototype(19 posts)→

Original post →

More from Models

Models channel →