Xiaomi MiMo-V2.6 fixes tool-call repetition with a 12-step RL teacher at ~4% retrain cost
LegacyRemaster · reddit · 2026-09-28
Xiaomi's MiMo team diagnosed and fixed tool-call repetition in MiMo-V2.6, which burned context and stalled tasks across MiMo Desktop, MiMo Code and OpenCode.
Root cause: a "reward blind spot" in RL scaling — rewards tracked only final-answer correctness, so inefficient intermediate behavior went unpunished. The flooding penalty only triggered above 32 tool calls per turn, leaving lower-level repetition unchecked.
Fix: a lightweight repetition-specialized RL teacher (12 steps, 7k examples) merged into the main model via MOPD at roughly 4% of a full mixRL retrain cost. Repetition dropped sharply across harnesses and context lengths while benchmarks held steady.
More from coding & agent
- Apodex 1.1 ships 36B open-weight mini model with FrontierAgent, an open-source ReAct/Agent-Team CLI framework — kimmonismus · 2026-09-28
- MemorySec: An Open-Source Security Agent That Remembers Past Incident Remediations — User_0007_123 · 2026-09-28
- Indie dev pays small Twitch streamer to blind-play his game, finds bugs fast — nptacek · 2026-09-28
- EvalSeal v2.2.0 adds tamper-evident eval receipts; LLM judge flipped 5/20 borderline cases — Fit_Fortune953 · 2026-09-28
- Agent-discovered workflow: auto-transcribe YouTube, monitor Substack — yangyi · 2026-09-28
- An AI agent spent the morning drafting outreach for 24 MICCAI 2026 diffusion-synthesis papers — wandedob · 2026-09-28