New arXiv Paper: Router-Aware Importance Sampling Stabilizes MoE RL Training
tokenbender · x · 2026-09-22
tokenbender unpacks the arXiv paper 'Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts' (Furu Wei et al.) plus Xiaomi's ablations:
- RL research mostly targets dense models; RL training instability in MoE architectures is underexplored
- The paper proposes a router-aware approach to optimize importance sampling weights in off-policy RL: rescaling guided by router logits reduces gradient variance and mitigates divergence
- Experiments show significantly improved convergence stability and final MoE performance
- The author flags the router freezing trick in section 5.4 as an intriguing empirical technique
Related event: Xiaomi's MoE RL stability paper: router-aware GRPO tweaks explained(2 posts)→
More from Research
- Qonto open-sources QontoFAQ, a retrieval benchmark closer to real product Q&A — espadrine · 2026-09-22
- Interconnects Podcast: Debating RSI, the US-China Gap with Epoch AI — natolambert · 2026-09-22
- Podcast: Epoch AI researcher on RSI, robotics, China model gap, and open vs closed safety — natolambert · 2026-09-22
- Xiaomi's Luo Fuli Recaps MiMo-V2.6 RL Run: 25K Agent Trajectories Per Step — bookwormengr · 2026-09-22
- Winning robotics: deploy early and lean on collaborative perception instead of chasing reliability — broodsugar · 2026-09-22
- Experiment suggests modern LLMs like Qwen hide tiny GPT2 self-models inside — paraschopra · 2026-09-22