Xiaomi Spent $2.6M on RL Post-Training and Lifted MiMo-V2.6's DeepSWE Score from 58.4 to 72.6

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Xiaomi LLM-Core Team, :, Zongming Qiao, Ziyue Hua, Zirui Ou, Zihao Yue, Zihan Jiang, Zhuo Huang, Zhiyang Chen, Zhixian Zheng, Zhipeng Xu, Zhengrui Ma, Yuyang Hu, Yuhang Dong, Yuechen Zhang, Yudong Wang, Yuanxin Liu, Yixin Yang, Yishuo Cai, Yikai Zhao, Yihan Yan, Yifan Zhang, Yifan Song, Xiyu Wei, Xing Zhang, Xin Zhang, Xiaoqian Liu, Xiaodong Ji, Xiangwei Deng, Xueyu Guo, Wenhan Ma, Weimin Xiong, Weikun Wang, Weiji Zhuang, Shuo Liu, Shuhuai Ren, Shuhao Gu, Shimao Chen, Shijie Cao, Shihua Yu, Shicheng Li, Shengjie Zhou, Shaolei Zhang, Rang Li, Qiying Wang, Qingkai Fang, Qianli Chen, Minzheng Wang, Liwen Wang, Linli Yao, Linghao Zhang, Liangyu Cheng, Liang Zhao, Lei Li, Jinhao Dong, Jinyu Xiang, Jianyu Wei, Jiangshan Duo, Huaqiu Liu, Huanjie Fan, Hongyi Guan, Hongshen Xu, Hao Tian, Hanyu Li, Hailin Zhang, Gang Wang, Fuli Luo, Feng Wei, Dong Zhang, Dawei Zhu, Chiheng Lou, Chenhong He, Chenhao He, Chenghua Liu, Bowen Ye, Bowen Shen, Boshen Xu, Bo Yang, Bingquan Xia, Bangjun Xiao, Baixuan Xu, Zhouxiang Mao, Zhiyang Zhang, Zhixiang Xu, Zhenru Lin, Zhengju Tang, Zhaojun Huang, Yuzhe Weng, Yuxing Xiang, Yuxiao Li, Yuheng Yang, Yuhang Wang, Yuchen Liu, Yuanyuan Tian, Yuanliang Dong, Yu Cheng, Yongzhe He, Yongshun Liang, Yong Wang, Yiyan Wang, Yitian Gong, Yijie Zhang, Yanshu Xin, Xun Zhang, Xingjian Zhao, Wenyu Yang, Wenshan Huang, Wenhao Li, Tingwei Huang, Tianyu Yu, Tianyang Lu, Taoyu Yang, Sinan Du, Shutong Tian, Shulin Du, Shengfan Wang, Shanchuan Fang, Qihao Zhang, Qibin Yang, Qian Yu, Qian Tu, Pengrong Xie, Peipei Wang, Peidian Li, Minkun Guo, Mingchen Shao, Luohan Gao, Lijie Wang, Liang Shi, Kaiqi Chen, Kaiming Liu, Kaifei Wang, Kai Yang, Jinlong Xue, Jiechen Zhang, Jiaxuan Liu, Hongxu An, Hao Peng, Hanglong Lü, Guonan Wang, Feiyu Yang, Fanyu Cao, Fangyue Liu, Fan Cui, Cong Wang, Chun Chen, Chenxu Bai, Chengxuan Zhu, Chenghua Wang, Boyi Zeng

cs.CL

2026-10-08

Xiaomi's MiMo-V2.6: $2.6M of agentic RL lifts a 1T-param MoE from 58.4 to 72.6 on DeepSWE, near Claude Opus 5's 74.0; RL environments, framework, and training logs are open-sourced.

What problem this solves

Scaling RL on long-horizon agent tasks is blocked less by the algorithm than by three pieces of unglamorous engineering: reward signals that are too coarse (a passed test is a 1, and says nothing about whether the solution is clean or lucky), agents that game the reward, and MoE routers that collapse under large-batch RL. Xiaomi's report rebuilds the whole stack for MiMo-V2.6-Pro (1.02T params, 42B active) and Flash (310B, 15B active). The previous MiMo-V2.5-Pro scored 19.0 on DeepSWE; Pro now reaches 71.9, 2.1 points behind Claude Opus 5. The stated ambition is recursive self-improvement, but the execution is concrete: one mixed-task RL run that puts coding, general, visual, and security tasks into the same training batch.

Method

Base models. The backbone interleaves sliding-window attention (window 128) with global attention, plus a 681M MiMo-ViT, audio encoders, and a DFlash-style MTP speculative decoder that drafts 7 tokens per pass. Pro pretrains on 30T tokens, Flash on 48T. An agent-centric mid-training stretches context from 256K to 1M and swaps AdamW for Muown, a Muon variant whose orthogonalized matrix updates hold up better at large batch; the switch caused no loss spikes.

RL is then scaled along three axes:

Router freezing is the stability key. With a trainable router, expert-load CV at one layer rose from 0.78 to 2.0 within the first 20 steps and cold experts went from 0.5% to 22%. Restoring just the step-20 router parameters to pre-RL values recovered the balance with no performance change, so the router drifted and the experts did not. Frozen, all three statistics stay flat.

After RL, MOPD2 distills from multiple domain teachers to cover hard-to-verify areas such as game development, research, and embodied intelligence.

Results

The RL run is itself a result: on DeepSWE v1.1 (average@3), Pro went from 58.4 to 72.6 for $2.6M and Flash from 48.7 to 65.7 for $0.9M. Pro's spending splits into 43.8% rollout, 43.5% training, and 12.7% grader.

BenchmarkMiMo-V2.6-ProMiMo-V2.5-ProClaude Opus 5GPT-5.6 Sol
DeepSWE v1.171.919.074.073.0
AutomationBench v1.0.653.116.050.345.8
CyberGym94.040.022.130.3
MiMo Visual Coding72.3—70.073.4
ProgramBench26.512.537.025.0

Three ablations, three answers. Without GAR, turn counts and token lengths explode and pass-rate gains stall; audits show the policy drifting toward speculative compatibility branches and swallowed exceptions, code that passes tests but is harder to maintain. With GAR, gains hold through step 52 and turns stay flat. Multi-harness training transfers: three harnesses never seen in training (claude code, codex, mini-swe-agent) improved mean Pass@1 from roughly 50% to 66% on DeepSWE. Anti-hacking defenses (environment cleanup, an iterative hack agent, offline audits during training) keep confirmed-hack trajectories below 2% throughout.

The failure analysis is unusually candid: Pro's 30 steps took 123.1 hours, interrupted by GPU double-bit errors, a Kubernetes crash, an unreachable grader, and an OOM caused by one expert receiving over 30× the mean token load in a micro-batch.

Why it matters

Half the value is the ledger. Few frontier-lab reports have itemized an agentic RL run this way: $2.6M total, grader at only 12.7%, every infrastructure failure logged. Transferable conclusions: freeze the MoE router before RL at this scale; use Muon-family optimizers for large batches; spend grader compute on within-group comparison to fight verbosity and reward hacking; build environment diversity from minimal recomposable harnesses. What is open-sourced is the RL stack itself, including verifiable task environments, the training framework, training dynamics, a 9B distilled model, and the mini-harness, a usable baseline for agentic RL teams.

Limitations

Terms

Source

Related papers

All paper explainers