Muon Optimizer's Spectral Space Trick
willdepue · x · 2026-07-14
The author shares a trick for the **Muon optimizer**: rather than hiding image information in the parameter's entry space, it's better to "hide it in the spectral space," as the optimizer updates more gently along its own norm directions. The original post explains that Muon diffuses updates across more gradient directions, making it better at erasing information written directly into the weights. However, it allocates less noise to fixed-direction updates, so encoding information into the spectral space allows it to survive training. The author also mentions that this method showed "very good retention" during the **NanoGPT speedrun**.
More from Research
- Draft paper uses Markov-chain eigenfunctions to build partitions and speed up sampling — michaelchchoi · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- GitHub repo adds lightweight ternary QAT for Prism-ML Bonsai models — terminoid_ · 2026-07-21
- Qdrant co-hosts a Munich meetup on search, retrieval, and agentic RAG on July 23 — qdrant_engine · 2026-07-21
- GigaChat Audio targets long-form audio grounding with timestamps across 120-minute inputs — ai-sage · 2026-07-21
- Paper models Transformer components as stochastic geometry and tests five architectures — Zhihua Liang · 2026-07-21