MiMo-V2.6 mid-RL run: ~2B tokens per step, multi-harness agentic RL, details to be open-sourced
bodonoghue85 · x · 2026-09-17
Xiaomi's MiMo team (@LuoFuli) says six months of silence went into studying how far RL can scale, and MiMo-V2.6 is now in the middle of its RL run, with training streamed live.
Three things were scaled:
- Compute: 2B tokens per step (1568 prompts × 16 rollouts), fully async
- Environments & harnesses: multi-task agentic RL, mixed across multiple harnesses in one run
- Grader compute: agentic in-group credit assignment with test-case and rubric-based rewards
Details will be open-sourced piece by piece over the coming weeks.
Related event: Xiaomi live-streams MiMo-V2.6 RL training with public dashboard(5 posts)→
More from coding & agent
- Baseten launches Hosted Tools, bringing server-side web search to open models — baseten · 2026-09-17
- Dev demos fast app testing with typesafe's Jev and opencode browser automation — soumitrashukla9 · 2026-09-17
- Demo shows blazing-fast browser automation with typesafe's Jev and opencode CLI — soumitrashukla9 · 2026-09-17
- 68-Agent Build Cut $2K in API Costs by Keeping Long Context on One Orchestrator Only — TheMoonMidas · 2026-09-17
- Dev proposes predictive dynamic context caching for Claude Code — Sauers_ · 2026-09-17
- Devin Fusion called best intelligence-per-cost coding agent; Meta Muse as wildcard pick — brandon_galang · 2026-09-17