Xiaomi's MiMo-V2.6 tech report: ~7k RL training data open-sourced, model shipped within a week of final RL run

rbhar90 · x · 2026-09-23

Xiaomi released the MiMo-V2.6 omni-modal tech report, centering on scaled RL toward self-improvement: asynchronous training consumes 1,568 samples and 2.73.7B tokens per step at up to 1M context; RL environments span code, general, visual, and cyber domains; groupwise agentic grading yields better reward signals, with a frozen MoE router and multi-layer defenses against reward hacking. The team plans to open-source 7k RL training samples plus the environments and framework, and shipped the model and report less than a week after starting the final RL run.

Original post →

More from Models

Models channel →