MiMo paper reveals 1.27M RL trajectories, but V4.1 Flash still looks better-baked
teortaxesTex · x · 2026-09-23
The author argues that despite benches and hype, MiMo-Flash is less thoroughly trained than V4.1 Flash, which got an absurd RL volume — millions of concurrent sandboxes. Citing MiMo's paper: 1.27M sampled trajectories (752K retained), one sandbox per rollout, rollouts taking 44% of the 82h run (36h for V2.6 Flash). They estimate one DSec unit plus V4.1 could match that rollout volume in 10 hours, and wonder what weeks more of RL would yield.
Related event: Community Estimates MiMo RL Training Scale, Still Short of V4.1(2 posts)→
More from Models
- Alibaba's Eddie Wu: Qwen sees RSI progress, plans 5-10T parameter model — teortaxesTex · 2026-09-23
- User Gives Opus 5.5 Creative Tools and Asks What It Dreams About — angrypenguinPNG · 2026-09-23
- Yuchen Jin: Opus 5.5 underwhelms, frontier LLM coding has plateaued — Yuchenj_UW · 2026-09-23
- Forward Future puts Opus 5.5 through 8 tests: cities, games, animation — MatthewBerman · 2026-09-23
- Matthew Berman: Opus 5.5 is the best model in the world — MatthewBerman · 2026-09-23
- A $5, 10-minute SFT run boosts Qwen3.6 by 8-12% on GPQA and MMLU-Pro — simonguozirui · 2026-09-23