Xiaomi's GAGAR: quality-aware advantage redistribution improves code agent RL
XiaomiMiMo · hf · 2026-09-29
Xiaomi's MiMo team introduces GAGAR, a framework fixing a core GRPO weakness in code-agent RL: all test-passing trajectories in a group get identical advantages, ignoring implementation quality.
Method
- Dynamic sampling keeps groups with both passing and failing trajectories
- An SFT-trained agentic grader inspects all group trajectories in a shared workspace and ranks passing candidates
- Lower-ranked trajectories are downweighted; advantages of passing trajectories are rescaled sum-preservingly to shift credit toward cleaner implementations
Results
- Validated at industrial scale on MiMo-V2.6-Flash (310B) and MiMo-V2.6-Pro (1.02T) pre-RL SFT checkpoints
- Controlled code-only Flash experiments: better code-agent performance, reduced trajectory-length growth, more stable training
- Also applied in large-scale mixed-task RL across Flash and Pro
Conclusion: combining test-based verification with groupwise agentic grading improves quality and stability of code agent RL.
More from coding & agent
- RSI Arena: 8 AI agents get 1,000 GPU-hours each to train a better model live — my_cat_can_code · 2026-09-29
- One-shot music video: Opus 5.5 + Runway MCP produces 'Words into Worlds' — tlakomy · 2026-09-29
- Anthropic breaks down Claude effort levels: when Max is overkill and when it pays off — rubenhassid · 2026-09-29
- MIT researcher: Opus 5.5 built an interactive JS slide deck for a 15-minute talk — davidbau · 2026-09-29
- Open-source VeriChron AI rebuilds past compliance states with bitemporal data, knowledge graphs and agents — LoksaiPalyam · 2026-09-29
- An ADR prompt for AI: how Staff+ engineers make architecture decisions — JafarNajafov · 2026-09-29