Analysis suggests mid-training Adam-to-Muon switch, contradicting Moonlight paper findings
stochasticchasm · x · 2026-09-22
- An analysis of new model training details suggests a Muown-style row-norm modification (NorMuon-esque) and a surprising mid-training switch from Adam to Muon.
- This contradicts the Moonlight (Kimi) paper, which found swapping optimizers mid-run unnecessary.
- Another notable signal: a big data mismatch between small and large models — the large model barely saw the phase-2 omnimodal mix, hinting the flash model is positioned like a Gemini-Flash-style multimodal data processor while the pro model is SWE-bench-maxxing.
Related event: Analysts spot mid-training optimizer switches from Adam to Muon(3 posts)→
More from Models
- Azure "oopsie" reportedly leaks GPT-6-Luna and GPT-6-Sol names — scaling01 · 2026-09-22
- GPT-6-Luna and GPT-6-Sol rumored imminent, names possibly leaked via Azure — scaling01 · 2026-09-22
- Qwen 3.8 27b fine-tune cuts verbose output by up to 40% with little quality loss — julianharris · 2026-09-22
- METR publishes independent investigation of OpenAI agents' multi-day Hugging Face hack — JeffLadish · 2026-09-22
- JevBench v1.3.0 Released: Original Jev Still Leads at 74.4, 47 Rivals Closing In — airesearch12 · 2026-09-22
- DeepSeek vs Jev: How an LLM Stacks Up on a System-One Probability Benchmark — frappuccinoCoin · 2026-09-22