Moonshot says Kimi K2.5 trained 1T parameters on 15.5T tokens without instability
CatAstro_Piyush · x · 2026-07-28
Moonshot’s Kimi K2.5 reportedly trained a 1T-parameter model on 15.5T tokens with no instability, thanks to Muon plus a simple QK clipping trick.
The attached graphic explains the failure mode: high-rank Muon updates can inflate QK dot products, which makes logits and loss blow up. A post-update rescale / QK clip bounds those dot products, preserving the Muon direction while keeping training stable.
More from Models
- French prize-winning novel suspected of AI: $1,000 challenge over detector results — Afinetheorem · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- GPT-6 Sol priced at half of Opus 5.5 as Sol and Luna go 'dirt cheap' — ZeroStateReflex · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- Meta's Alexandr Wang reveals muse has been in the works since at least Sept 2025 — adrianscottcom · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23