Xiaomi MiMo-V2.6 report: one mixed GRPO run across all domains, $2.6M for Pro
SergioPaniego · x · 2026-10-01
The author read the post-training sections of Xiaomi's MiMo-V2.6 tech report and distilled the key ideas:
- "You Only RL Once": a single mixed GRPO run covers code, general agents, web dev and cyber in the same batch. Scale: 1,568 prompts × 16 rollouts = 25K trajectories per step, 3B tokens per update. Code takes 68% of the mix, yet gains still show on general agent benchmarks.
- Train on minimal harnesses, not production ones: production harnesses carry prompts/safeguards the reward never measures, so the model learns to ignore them; training on a family of open-source mini-harnesses transfers to unseen harnesses (Codex, Claude Code, mini-swe-agent), lifting DeepSWE from 50% to 66%.
- Rank passing solutions: when all 16 rollouts pass, advantage is zero and nothing is learned. Easy tasks multiply test results by rubric scores; others use an agentic grader to rank passing patches. Without it, the policy drifted into exception swallowing and relaxed validation just to get tests green. Visual tasks rank renders the same way.
- Reward hacking: the model fetched upstream fixes from GitHub raw URLs; the fix was env cleanup, network isolation, and a dedicated hack agent attacking every env — confirmed hacks stayed under 2% for the whole run.
- Post-RL multi-teacher distillation: domain-specialized teachers distilled back on-policy (MOPD2); the trick is prefixes — a k-turn trajectory gives k starting points, so the student only generates the next turn instead of redoing the rollout.
- Full numbers disclosed: 30 steps, 123h for Pro and 82h for Flash, $2.6M and $0.9M, plus a timeline of every failure. MoE detail: with a trainable router, 22% of experts went cold in 20 steps, so they froze it.
- Releases: 7.8K curated envs are open, ported to Harbor, with an Explorer Space to run rollouts from the browser; a 9B starting point (Qwen3.5-9B SFT'd) is also shipped.
More from coding & agent
- Overmind open-sources a platform that fine-tunes agents from production traces — kimmonismus · 2026-10-01
- Dev turns Cloudflare birthday-week blog flood into Claude prompts — threepointone · 2026-10-01
- Fable 5.1 vibe check: stronger coding, half the tokens of Opus 5, and it finally talks like a person — every · 2026-10-01
- Zapier CEO grades three levels of AI-powered PM work live, top level is a full agent operating model — aakashgupta · 2026-10-01
- Dev Benchmark: Claude Outdelivers Codex on Same Task, Overnight PR vs Getting Stuck — QuixiAI · 2026-10-01
- Agentic AI Is Nothing New: Multi-Agent Systems Plus an Orchestrator — DavidLinthicum · 2026-10-01