Kimi K3 trains its vision encoder from scratch and claims better vision evals
nrehiew_ · x · 2026-07-29
The thread says Kimi K3 is natively multimodal and that its vision encoder is trained from scratch.
- Instead of training the encoder first or initializing from a contrastively pretrained model, K3 trains the whole system jointly with next-token prediction.
- The author says this approach beats baselines on vision evaluations.
- The attached figure contrasts Kimi K2 and K3 scaling curves and includes a plot showing the new vision tower, MoonViT-V2, with lower gradient norms and fewer spikes than the SigLIP-initialized MoonViT-3D.
- The post argues that contrastive pretraining may be unnecessary as a multimodal initialization at scale.
Related event: Deep Dive into Kimi K3 Tech Report: Architecture and Training(18 posts)→
More from Models
- Grok 4.5 Medium tops LaurenBench with 56.9%, ahead of Claude Sonnet 5 and GLM 5.2 — elonmusk · 2026-07-29
- A viral chart compares 12 paid AI tools with free replacements — nikola_mr64990 · 2026-07-29
- Verdent Partners with Moonshot to Optimize Agentic Coding for 2.8T-param Kimi K3 — PrajwalTomar_ · 2026-07-29
- Best Local Models Under 120B: Are Qwen Series the Only Answer? — Possible_Grocery8079 · 2026-07-29
- OpenMed Launches Open-Source Medical AI Framework with Local Kimi Integration — MaziyarPanahi · 2026-07-29
- Professor Tests Flux 3: Extremely High Prompt Adherence — emollick · 2026-07-29