GLM 5 follows up with per-head Muon to balance attention-head updates
stochasticchasm · x · 2026-07-28
- The post highlights a paper detail about per-head Muon following GLM 5.
- The image explains that for attention projections, the optimizer is refined into a per-head variant rather than applying Newton–Schulz orthogonalization to the full Q/K/V matrices.
- The motivation is that full-matrix orthogonalization can let large-gradient heads dominate the shared update, while per-head orthogonalization balances learning across heads.
- In practice, the approach reportedly improves stability at larger scales and slightly reduces optimizer overhead because Newton–Schulz iterations on smaller per-head blocks are cheaper.
More from Research
- CMU’s O-VAD detects industrial video anomalies by tracking object state over time — CarnegieMellonU · 2026-07-28
- Latent Action Model talk will show how to learn Super Mario from observation alone — ceciletamura · 2026-07-28
- MLA-based KV cache costs 12 GB per million tokens, with KDA state at 230 MB BF16 — zephyr_z9 · 2026-07-28
- OpenAI chart says 43.5% of occupation-specific ChatGPT use goes beyond the user’s job — soumitrashukla9 · 2026-07-28
- OpenAI economic research says AI is reshuffling tasks across occupations — soumitrashukla9 · 2026-07-28
- AI brute-forces Erdős problems: Pushing multiple math frontiers in a single day — kevrussell · 2026-07-28