Muon-trained models leave a distinctive spectral signature in their weights

QuanquanGu · x · 2026-09-14

Research shows you can tell whether a model was trained with the Muon optimizer just by examining the empirical spectral distribution of its hidden-layer weights.

Compared to SGD/Adam, Muon leaves a highly distinctive signature in W^T W, because it updates low-magnitude directions more aggressively. This means optimizer choices leave a detectable fingerprint, with implications for model provenance and reproducibility research.

Original post →

More from Research

Research channel →