New paper explains sudden learning & scaling laws via "permutation symmetry"
CatAstro_Piyush · x · 2026-08-18
A new preprint (arXiv:2608.13335) attempts to explain why all neural networks exhibit phenomena of "sudden learning" and follow "scaling laws". The study identifies permutation symmetry as the root cause.
Key insights:
- At small initialization, permutation symmetry restricts training dynamics to depend solely on an architecture-specific "structure matrix A".
- Perceptrons, attention layers, MoE, and convolutions all reduce to this minimal model with different A matrices.
- Dynamics close on the "order parameter" M=WW^T and reduce to a Lotka-Volterra equation where modes switch on sequentially.
- Smaller initial weights lead to wider gaps between switch-on times, causing loss plateaus followed by sudden drops; unresolved modes merge into a smooth power law scaling.
More from Research
- Yoshua Bengio's classic 'Deep Learning' textbook is free to read online — JafarNajafov · 2026-08-24
- UnsolvedMath Dataset Grows to 8,785 Problems, Features GPT-5.6 Generated Solutions — lhoestq · 2026-08-24
- Weekend project: Simulating fruit fly connectome to drive virtual rover — anselm · 2026-08-24
- TACL Journal Seeks Reviewers with PhD and High Citation Counts — EhudReiter · 2026-08-24
- New Papers on Superposition Spark 2019 Nostalgia Among Researchers — _xjdr · 2026-08-24
- Biomni: A General-Purpose AI Agent for Executing Biomedical Research Workflows — bravo_abad · 2026-08-24