WeightWatcher experiments suggest Muon beats AdamW by reducing memorization
rasbt · x · 2026-09-19
The WeightWatcher team shared new experiments on the AdamW vs Muon debate, suggesting Muon's edge may come from reduced memorization. They trained a single-head nanoGPT model with both optimizers across five seeds, then analyzed weight-matrix spectra and several forms of memorization.
Key observations: AdamW yields relatively stable spectral α values, while Muon is much noisier with layer-wise ESDs varying substantially, making α harder to estimate. Crucially, individual layers under Muon drop below α = 2 — a regime their spectral/RG theory associates with pathological, example-specific memorization. The layer-level detail hidden by averages may explain the generalization gap.
More from Research
- PAW edges out Gemma e4b in early tests, with reliability as the real win — yuntiandeng · 2026-09-19
- EPFL and Oxford researchers launch call for open science in AI safety, petition frontier labs — manoelribeiro · 2026-09-19
- Hamel Husain Explains What a Trace Is in LLM Evaluation — HamelHusain · 2026-09-19
- Forget model names, look at the ops: the classic math behind attention, RoPE and diffusion — techNmak · 2026-09-19
- Researcher uses AI to illustrate how AI writing is breaking peer review — sethlazar · 2026-09-19
- Who are the poor souls chairing ICLR 2027? Time to rethink publishing norms — CharlotteHase · 2026-09-19