Microsoft's ESRL boosts MoE RL via expert-space exploration
MicrosoftResearch · hf · 2026-09-15
- Microsoft Research proposes ESRL to improve reinforcement learning for mixture-of-experts language models.
- Key techniques: exploring expert routing anchored on high-confidence experts, entropy-adaptive perturbation, and path replay to increase rollout diversity.
- The goal is better exploration and training performance for MoE models under RL.
More from Research
- Grouped Value Attention shrinks KV cache by reconstructing keys on demand — Vishesh Tripathi · 2026-09-15
- Amazon's MInTRL uses sparse off-policy interventions to boost on-policy RL — amazon · 2026-09-15
- Stateless LLM failover preserves ~0% context; ContinuityBench proxy hits 99.20% CPR — its_vayishu · 2026-09-15
- Phillip Isola highlights a non-mainstream AI route: RL from scratch via ultra-fast simulators — AjdDavison · 2026-09-15
- Cutting AI verifier reading cost: top-50 retrieval kept just 2 of 8 minority evidence items — iMiguelmars · 2026-09-15
- SSAD2026 talk covers autonomous driving 3D perception, from LiDAR self-supervision to multi-sensor distillation — abursuc · 2026-09-15