Classic MiMo architecture with huge RL batches: 1568 tasks x 16 rollouts per step, surprising results
stochasticchasm · x · 2026-09-22
The author notes MiMo uses a classic architecture — simple yet strong. The RL run used an enormous batch size at every step: 1568 tasks with 16 rollouts each, meaning the run consumed a lot of tokens. Even so, the results remain surprisingly good.
More from Research
- Paper author: test-time communication turns parallel search into cumulative discovery — DimitrisPapail · 2026-09-22
- Berkeley paper: communicating agent teams match 4x more independent agents on ARC-AGI-3 — DimitrisPapail · 2026-09-22
- Lean vs ZFC: the rules of mathematical proof weren't changed by any vote — jessi_cata · 2026-09-22
- David Krueger: four unresolved foundational problems stand between us and safe AI — DavidSKrueger · 2026-09-22
- Multi-agent scaling can be compute-optimal: N parallel agents beat one agent run N-times longer — DimitrisPapail · 2026-09-22
- RL training config debate: 30 steps x 25k rollouts is wild, steps ≈ rollouts is the sane default — willcb · 2026-09-22