MiMo-V2.6 mid-RL run: ~2B tokens per step, multi-harness agentic RL, details to be open-sourced

bodonoghue85 · x · 2026-09-17

Xiaomi's MiMo team (@LuoFuli) says six months of silence went into studying how far RL can scale, and MiMo-V2.6 is now in the middle of its RL run, with training streamed live.

Three things were scaled:

Details will be open-sourced piece by piece over the coming weeks.

Related event: Xiaomi live-streams MiMo-V2.6 RL training with public dashboard(5 posts)→

Original post →

More from coding & agent

coding & agent channel →