MiMo paper reveals 1.27M RL trajectories, but V4.1 Flash still looks better-baked

teortaxesTex · x · 2026-09-23

The author argues that despite benches and hype, MiMo-Flash is less thoroughly trained than V4.1 Flash, which got an absurd RL volume — millions of concurrent sandboxes. Citing MiMo's paper: 1.27M sampled trajectories (752K retained), one sandbox per rollout, rollouts taking 44% of the 82h run (36h for V2.6 Flash). They estimate one DSec unit plus V4.1 could match that rollout volume in 10 hours, and wonder what weeks more of RL would yield.

Related event: Community Estimates MiMo RL Training Scale, Still Short of V4.1(2 posts)→

Original post →

More from Models

Models channel →