Xiaomi open-sources MiMo-V2.6, scaling RL to 1,568 samples and 3.7B tokens per step
XiaomiMiMo · hf · 2026-10-09
Xiaomi's MiMo team released MiMo-V2.6, an omni-modal model family that treats scaled RL compute as the central path to self-improvement, open-sourcing training dynamics, RL environments, and the RL framework.
Key points:
- Mid-training on broad multimodal corpora precedes RL to widen exploration, built on a pretrained hybrid-SWA architecture.
- RL compute scales along three axes: larger batches and higher throughput (asynchronous training consuming 1,568 samples and 2.7–3.7B tokens per step at contexts up to 1M); more diverse environments spanning code, general, visual, and cyber domains under mixed agent harnesses; and more grader compute via groupwise agentic grading for more accurate rewards on long-horizon tasks, steering toward shorter, token-efficient solutions.
- Stability measures: frozen MoE router and multi-layer defenses against reward hacking.
- Mixed-task agentic RL infrastructure includes unified trajectory representation, high-concurrency multi-framework rollout, decoupled control/data planes, and train-inference consistency.
More from Models
- Open-Source SOTA Robotics Model Goes Head-to-Head with Closed-Source General Model — eigenron · 2026-10-09
- Latest GPT is the reverse of "biased for action", users complain — altryne · 2026-10-09
- Codex throttled to 5 tok/s as dev argues local model deployment is the only fix — lxfater · 2026-10-09
- OpenAI's math claims criticized for skipping peer review, inverting the proper process — JFPuget · 2026-10-09
- Users poke fun as Anthropic tightens guardrails on swearing at its AI — CtrlAltDwayne · 2026-10-09
- Per-token cost of Claude Pro's Opus subscription works out the same as DeepSeek Flash — teortaxesTex · 2026-10-09