Xiaomi MiMo-V2.6 scales RL to 3.7B tokens per step, targeting self-improvement
Dr_Singularity · x · 2026-09-23
- Xiaomi's MiMo-V2.6 is its biggest RL scaling effort yet, spanning coding, general reasoning, vision, automation, and cybersecurity.
- Each RL step processes 1,568 samples and 2.7–3.7 billion tokens, with context lengths up to 1M tokens.
- Instead of narrow benchmarks, training uses multiple complex environments and stronger agentic graders to improve long-horizon reasoning and push for more efficient solutions.
- Benchmark performance keeps rising throughout training, suggesting large-scale RL still has substantial headroom; Xiaomi explicitly frames this as scaling toward model self-improvement.
More from Models
- Users marvel at Opus 5.5 being 'too good to be true', brace for a nerf — iamsahaj_xyz · 2026-09-23
- Claude Opus 5.5 shown handling a code review, 'taking Theo's job' for the day — 0xkarasy · 2026-09-23
- Sarvam's Saaras V4 adds keyterm prompting to boost speech transcription accuracy — cneuralnetwork · 2026-09-23
- Xiaomi's MiMo V2.6 Pro tops open-source leaderboard at ~$0.13 per task — heyshrutimishra · 2026-09-23
- Xiaomi MiMo V2.6 Pro tops open-source leaderboard at 46, costs ~$0.13 per task — heyshrutimishra · 2026-09-23
- Deep-scan reveals Muse agent's hidden powers: Instagram DMs, HomeKit control, its own email identity — flaneur451 · 2026-09-23