Sticking with GRPO: The Four Targeted Mods Behind Xiaomi's Stable MoE RL
tokenbender · x · 2026-09-22
tokenbender breaks down Xiaomi's MoE RL recipe: no fancy algorithm swap, just targeted GRPO modifications for stability and scale:
- Prompt-mean loss instead of token-mean to stop long answers from dominating gradients (from DAPO)
- Separate clipping bounds for positive and negative advantages (DAPO)
- Entropy-adaptive clipping (mai-thinking-1)
- Segment-level credit shaping (likely TreeAdv)
Related event: Xiaomi's MoE RL stability paper: router-aware GRPO tweaks explained(2 posts)→
More from Research
- Qonto open-sources QontoFAQ, a retrieval benchmark closer to real product Q&A — espadrine · 2026-09-22
- Interconnects Podcast: Debating RSI, the US-China Gap with Epoch AI — natolambert · 2026-09-22
- Podcast: Epoch AI researcher on RSI, robotics, China model gap, and open vs closed safety — natolambert · 2026-09-22
- Xiaomi's Luo Fuli Recaps MiMo-V2.6 RL Run: 25K Agent Trajectories Per Step — bookwormengr · 2026-09-22
- Winning robotics: deploy early and lean on collaborative perception instead of chasing reliability — broodsugar · 2026-09-22
- Experiment suggests modern LLMs like Qwen hide tiny GPT2 self-models inside — paraschopra · 2026-09-22