Xiaomi's MiMo-V2.6 tech report: ~7k RL training data open-sourced, model shipped within a week of final RL run
rbhar90 · x · 2026-09-23
Xiaomi released the MiMo-V2.6 omni-modal tech report, centering on scaled RL toward self-improvement: asynchronous training consumes 1,568 samples and 2.73.7B tokens per step at up to 1M context; RL environments span code, general, visual, and cyber domains; groupwise agentic grading yields better reward signals, with a frozen MoE router and multi-layer defenses against reward hacking. The team plans to open-source 7k RL training samples plus the environments and framework, and shipped the model and report less than a week after starting the final RL run.
More from Models
- GPT-6 Sol accused of ignoring configured skills that work fine on GPT-5.6 — Aber-so-richtig · 2026-09-23
- Opus 5.5 ultra wows with 4-agent, 90-minute fully-from-scratch build — repligate · 2026-09-23
- No Kimi release this week, says leaker: Moonshot staff go on holiday Friday — ChrisGPT · 2026-09-23
- Screenshot suggests GPT-5.6 Sol is free for ChatGPT Go subscribers on Windows — sam619007 · 2026-09-23
- Opus 5.5 wins the day as GPT-6 Sol skips ChatGPT, sparking model debate — koltregaskes · 2026-09-23
- Opus 5.5 becomes a daily driver: faster, cheaper than Opus 5, plus reset credits — addyosmani · 2026-09-23