Xiaomi Open-Sources MiMo-V2.6, Scaling RL for Self-Improvement
Xiaomi's MiMo team released and open-sourced the omni-modal MiMo-V2.6 series, positioning RL compute as the central paradigm for self-improvement. The highly automated pipeline, trained with roughly $2.6M of RL compute and 1568 samples per step, pushed DeepSWE to 72.6.
2026-10-09 ~ 2026-10-10 · 2 related posts
- Xiaomi open-sources MiMo-V2.6, scaling RL to 1,568 samples and 3.7B tokens per step — XiaomiMiMo · 2026-10-09
- Xiaomi's MiMo-V2.6: agents run their own RL loop, DeepSWE score hits 72.6 for $2.6M — rohanpaul_ai · 2026-10-10