MoonViT-V2 trains Kimi K3’s vision encoder from scratch with next-token prediction
stochasticchasm · x · 2026-07-28
The image and reply discuss MoonViT-V2, a vision encoder trained from scratch with next-token prediction.
- The key claim is that Kimi K3’s vision encoder departs from Kimi K2.5 by not starting from a contrastively pre-trained model such as SigLIP.
- The paper argues that pre-trained encoders can make joint optimization unstable, while MoonViT-V2 shows lower gradient norms and fewer spikes during training.
- Next-token prediction is said to let the encoder’s representations be shaped directly by the language-model objective.
- The figure claims MoonViT-V2 matches the SigLIP-initialized baseline on vision evaluations, suggesting contrastive pre-training may be unnecessary at scale for multimodal language models.
Related event: Kimi K3 Report: SiTU-GLU and Stability in Large-Scale MoE Training(6 posts)→
More from Multimodal
- vLLM ships day-0 serving support for Moonshot’s 2.8T-parameter Kimi K3 — vllm_project · 2026-07-28
- Invideo’s Agent One turns a simple conversation into a cinematic trailer — LudovicCreator · 2026-07-28
- Invideo Agent One turns a conversation into a cinematic trailer end to end — LudovicCreator · 2026-07-28
- Midjourney 8.2 prompt generates a black-and-white Zendaya fashion shot in New York — michaelrabone · 2026-07-28
- Higgsfield opens up the full production blueprint behind its Originals — mhdfaran · 2026-07-28
- A Civitai image-edit model fails its own edit claim in a quick test — cradledust · 2026-07-28