Xiaomi MiMo-V2.6 fixes tool-call repetition with a 12-step RL teacher at ~4% retrain cost

LegacyRemaster · reddit · 2026-09-28

Xiaomi's MiMo team diagnosed and fixed tool-call repetition in MiMo-V2.6, which burned context and stalled tasks across MiMo Desktop, MiMo Code and OpenCode.

Root cause: a "reward blind spot" in RL scaling — rewards tracked only final-answer correctness, so inefficient intermediate behavior went unpunished. The flooding penalty only triggered above 32 tool calls per turn, leaving lower-level repetition unchecked.

Fix: a lightweight repetition-specialized RL teacher (12 steps, 7k examples) merged into the main model via MOPD at roughly 4% of a full mixRL retrain cost. Repetition dropped sharply across harnesses and context lengths while benchmarks held steady.

Related event: Xiaomi MiMo-V2.6 Fixes Repeated Tool-Call Bug Caused by RL Reward Blind Spot(2 posts)→

Original post →

More from coding & agent

coding & agent channel →