Manifold Muon offers a loss-free path for training MoE routers
tokenbender · x · 2026-07-25
A post from Varun Neal describes two methods for training MoE routers using Manifold Muon, including one approach that is completely detached from the training loss.
The author says the detached version, when fed only an explicitly calculated load-balancing gradient, already works well on its own, though it still trails the current state of the art. The shared code snippet shows a simple router step that combines a balancing gradient with a Manifold Muon update.
More from Research
- Zvi says Opus 5 matches Fable on virology tasks, raising a safety-policy question — TheZvi · 2026-07-25
- Opus 5 appears to improve on ARC-AGI 1 and 2, and may rely on algebraic puzzle solving — herbiebradley · 2026-07-25
- NSF funds network of AI-enabled cloud labs for autonomous science — ProfBuehlerMIT · 2026-07-25
- Blocked biology requests on Fable 5 will route to Opus 5, Anthropic says — nc_frey · 2026-07-25
- AlphaFold-guided redesign cuts off-target risk in gene-editing proteins — Ars Technica AI · 2026-07-25
- Practical multi-agent orchestration for Codex splits work into scout, worker, and coordinator roles — pvncher · 2026-07-25