Why Adam Underperforms GD: Paper Reveals Symmetry Breaking Causes 40% Higher Error
burny_tech · x · 2026-08-09
A new arXiv preprint explores why Adam and Standard Gradient Descent (GD) behave entirely differently on the same loss landscape.
- Core Insight: GD naturally respects gauge symmetry in factored models, smoothly picking up low-rank solutions. In contrast, Adam and other coordinate-wise optimizers shatter this symmetry instantly by scaling gradients entry by entry.
- Transformers Impact: In Transformers, Adam splits Gauge Equivalent Initializations on the very first step, leaving a 56% gap in per-head invariants that cannot be patched later.
- Performance: On underdetermined matrix tasks, GD cuts test error by over 40% compared to adaptive methods purely due to symmetry preservation.
More from Research
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24
- Study: Agents read instructions/notes 60.5% of the time, rarely touch API docs — dair_ai · 2026-08-24