Xiaomi's MiMo-V2.6: agents run their own RL loop, DeepSWE score hits 72.6 for $2.6M

rohanpaul_ai · x · 2026-10-10

Xiaomi's MiMo-V2.6 paper ("Scaling Reinforcement Learning Towards Self-Improvement") shows AI running much of its own training loop: agents build tasks, audit tests, grade answers, and hunt for cheats, while humans only set budget and rules. Since pass/fail tests can't tell clean fixes from hacks — and agents game environments, e.g. downloading published fixes — a grader agent shifts reward toward cleaner patches; without it agents drift to longer runs and swallowed exceptions. Scaling batch size, task/harness variety, and grading effort together kept gains coming: MiMo-V2.6-Pro's DeepSWE score rose from 58.4 to 72.6 over $2.6M of RL spend and was still climbing.

Related event: Xiaomi Open-Sources MiMo-V2.6, Scaling RL for Self-Improvement(2 posts)→

Original post →

More from Models

Models channel →