Muon orthogonalization goes peer-to-peer to cut all-gather overhead
stochasticchasm · x · 2026-07-28
The post points to a paper section on P2P-based Muon orthogonalization.
- In distributed training, the optimizer shards parameters evenly across data-parallel ranks, but Muon’s Newtown–Schulz orthogonalization needs the full parameter matrix before each update.
- A naive all-gather on every rank creates a large memory footprint and makes communication the bottleneck at scale.
- The proposed approach lets each rank retrieve only the shards it locally owns via peer-to-peer communication with the corresponding owner ranks.
- This removes the full-parameter buffer, reduces communication volume, and pipelines communication with computation at the granularity of model-chunk buffers.
More from Infra
- Underlayer Electrons Aggravate Stochastic Defectivity in EUV Lithography — CatAstro_Piyush · 2026-07-28
- SkyPilot says serving Kimi K3 needs multi-node inference and a full stack — skypilot_org · 2026-07-28
- Ilya Sutskever’s SSI partners with Nvidia in a deal Bloomberg pegs at $5 billion — RebeccaBellan · 2026-07-28
- Moonshot’s Mooncake serving stack boosts long-context throughput by up to 525% — stochasticchasm · 2026-07-28
- Fireworks says Kimi K3 matches Opus 5 quality at 2x–4.6x lower task cost — lqiao · 2026-07-28
- MoE training paper proposes workload-aware expert-GEMM scheduling — stochasticchasm · 2026-07-28