Kimi K2.5 Scaling Methodology
AI寒武纪 · wechat · 2026-07-18
This article recaps Moonshot AI CEO Yang Zhilin's GTC speech on the "scaling methodology" behind Kimi K2.5. The main theme divides model capability scaling into three dimensions: **token efficiency, long context, and agent quantity**. - **Token efficiency**: The team detailed the Muon optimizer and its distributed implementation, claiming it can replace Adam and yield better training results under identical token/parameter constraints. They also used **QKclip** to resolve attention logit explosion and loss divergence during large-scale training. - **Long context**: They proposed **KimiLinear / KimiDeltaAttention**, mixing linear and full attention, and changed the original scalar decay to a diagonal matrix, balancing performance and efficiency across short-context, long-input, and long-output tasks. - **Agent swarm**: Designed an orchestrator + sub-agent training paradigm, paired with instantiation rewards, completion rewards, and outcome rewards, aiming to enable multiple agents to collaborate on complex tasks in parallel. The article also highlighted key aspects of K2.5: - Trained on over **15 trillion tokens** on an H800 cluster; - The first open-source model trained jointly on **native vision + text**, emphasizing that "early fusion" is superior to text-before-vision; - Visual and textual capabilities mutually enhance each other, achieving strong performance even with almost no visual SFT; - The speech concluded with a teaser for the next-gen architecture direction: analogizing "residual connections" to a temporal recurrent structure, and further evolving it into "attention residuals."
Related event: Yang Zhilin Shares Kimi K2.5 Scaling and Efficiency Strategies(4 posts)→
More from Models
- Epoch AI Live Streams GPT-5.6 Playing Slay the Spire — dr_cintas · 2026-07-21
- Side-by-side model test lands both answers on the first try, then shifts to loop engineering — glenbeer · 2026-07-21
- Emad Mostaque says Kimi K3 inference costs could fall 10x to 50x soon — rohanpaul_ai · 2026-07-21
- User shares a striking GPT-5.6 embodiment passage from a simple writing prompt — Comfortable_Ebb5519 · 2026-07-21
- Reddit asks whether Kimi K3 is already good enough for production agents — CommercialClient2408 · 2026-07-21
- Korean startup says its model scored 44 on AAII and matches DeepSeek V4 Pro — JungWooHa2 · 2026-07-21