Xiaomi's Luo Fuli Recaps MiMo-V2.6 RL Run: 25K Agent Trajectories Per Step

bookwormengr · x · 2026-09-22

Xiaomi's Luo Fuli published a long retrospective on MiMo-V2.6, framed around one question: pretraining scaling laws have run for years — can RL scale the same way, with more compute, more tasks and more real environments still translating into capability gains? Xiaomi has spent the past half year betting on it.

RL scale

Challenge 1: enough "worlds"

Math RL is one problem, one answer, instant feedback. Agents must write code, open browsers, operate computers, call tools and read results before deciding the next move. Xiaomi expanded its environments accordingly, building 7,000+ RL task environments spanning coding, general agent, visual agent and cybersecurity tasks, with trajectories lasting dozens to hundreds of steps.

Challenge 2: Harness enters training

Previously the harness lived outside the model — model makers trained the model, and products like Claude Code and Codex wrapped a system prompt around it. Xiaomi brought the harness into the training loop.

(The original post is a long thread; later sections were not fully provided.)

Original post →

More from coding & agent

coding & agent channel →