Nora optimizer keeps Muon's matrix structure benefits without the full cost
burkov · x · 2026-09-03
Burkov breaks down Nora, a new optimizer from teams at Nankai, Renmin, Tsinghua and XJTU:
- Problem: Training large transformers is partly an optimizer-design problem. Methods like Muon exploit the matrix structure of weights for better update directions, but that's computationally expensive; cheaper approximations can disturb weight norms — problematic in networks where scaling a weight changes little about its function.
- Approach: Nora first removes the component of momentum pointing along each row of the weight matrix, so updates change row direction rather than length, then normalizes each row.
- Result: This simple operation approximates a structured rescaling of the update (preconditioning), keeping the useful part of Muon-style structure at much lower cost.
More from Research
- Linear transforms beat deep learning (ScanVI) for single-cell batch correction, preprint shows — arjunrajlab · 2026-09-03
- SimLoss Enables Single-Pass Fine-Grained Image Captioning at Multi-Stage Quality — Suryaansh Jain · 2026-09-03
- Tenstorrent and AI & Inc launch JapanFold: free inference for open-source drug discovery models — DavidBennett__ · 2026-09-03
- Until Labs Scales Cryoprotectant Search to 250,000 Molecules With AI — NirantK · 2026-09-03
- X Debate: Is Chain-of-Thought Prompting a Form of Parameter Reuse? — aryaman2020 · 2026-09-03
- Constraining agents with LL(1) grammar + structured diagnostics: what it fixes and what slips through — Upstairs-Special-925 · 2026-09-03