Moonshot Details Kimi Training Path

dotey · x · 2026-07-18

In his GTC 2026 talk How We Scaled Kimi K2.5, Yang Zhilin detailed Moonshot AI's technical roadmap over the past year: how to keep approaching closed-source frontiers using open-source models.

Three core strategies:

He mentioned two specific advances: visual training can inversely enhance text capabilities, and the team just announced their next-gen architecture, Attention Residue. For foundational training components, Moonshot AI has attempted to replace the long-standing Adam optimizer, attention mechanisms, and residual connections, all of which are open-sourced.

Notably, MuonClip replaces Adam. Moonshot claims that for the same amount of data, its training effectiveness nearly doubles the data volume; as high-quality data becomes scarcer, this means higher data efficiency and lower training costs.

Related event: Yang Zhilin Shares Kimi K2.5 Scaling and Efficiency Strategies(4 posts)→

Original post →

More from Models

Models channel →