Moonshot says Kimi K2.5 trained 1T parameters on 15.5T tokens without instability
CatAstro_Piyush · x · 2026-07-28
Moonshot’s Kimi K2.5 reportedly trained a 1T-parameter model on 15.5T tokens with no instability, thanks to Muon plus a simple QK clipping trick.
The attached graphic explains the failure mode: high-rank Muon updates can inflate QK dot products, which makes logits and loss blow up. A post-update rescale / QK clip bounds those dot products, preserving the Muon direction while keeping training stable.
More from Models
- Users say Opus 5 breaks things, while Opus 4.8 still feels strong — omarsar0 · 2026-07-28
- Kimi K3 Max tops Frontend Code Arena overall and leads 5 of 7 domains — eyishazyer · 2026-07-28
- Kimi report reveals a wide internal benchmark suite for coding and agent skills — stochasticchasm · 2026-07-28
- Claude is still being called the most steerable model set, despite its weirdness — sloppenheimer · 2026-07-28
- Frontend Code Arena: Opus 5 Max Takes #1, Kimi K3 Max Follows Closely — arena · 2026-07-28
- Kimi K3 Max Tops Arena Leaderboard in Frontend Code and Agent Tasks — arena · 2026-07-28