Manifold Muon offers a loss-free path for training MoE routers

tokenbender · x · 2026-07-25

A post from Varun Neal describes two methods for training MoE routers using Manifold Muon, including one approach that is completely detached from the training loss.

The author says the detached version, when fed only an explicitly calculated load-balancing gradient, already works well on its own, though it still trails the current state of the art. The shared code snippet shows a simple router step that combines a balancing gradient with a Manifold Muon update.

Original post →

More from Research

Research channel →