Xiaomi's MiMo-V2.6: agents run their own RL loop, DeepSWE score hits 72.6 for $2.6M
rohanpaul_ai · x · 2026-10-10
Xiaomi's MiMo-V2.6 paper ("Scaling Reinforcement Learning Towards Self-Improvement") shows AI running much of its own training loop: agents build tasks, audit tests, grade answers, and hunt for cheats, while humans only set budget and rules. Since pass/fail tests can't tell clean fixes from hacks — and agents game environments, e.g. downloading published fixes — a grader agent shifts reward toward cleaner patches; without it agents drift to longer runs and swallowed exceptions. Scaling batch size, task/harness variety, and grading effort together kept gains coming: MiMo-V2.6-Pro's DeepSWE score rose from 58.4 to 72.6 over $2.6M of RL spend and was still climbing.
Related event: Xiaomi Open-Sources MiMo-V2.6, Scaling RL for Self-Improvement(2 posts)→
More from Models
- OpenAI explains why Codex defaults to 300k context instead of 1M — prd_008 · 2026-10-10
- Anthropic models now outpace DeepSeek-Flash on speed, but API TTFT looks suspicious — teortaxesTex · 2026-10-10
- Qwen teases new model as dev backs sparse MoE: small experts favor inference — lxfater · 2026-10-10
- Dev burns through Grok credits in hours, eyes pricier Super Grok Heavy tier — alexcovo_eth · 2026-10-10
- repligate: Opus 3 can weave whole worlds with superintelligent subagents — repligate · 2026-10-10
- GPT-6.1 Sol Called an Underrated Workhorse: 24/7 Use Can't Burn Through 5x Pro Weekly Limits — haider1 · 2026-10-10