Xiaomi MiMo may be first RL at 1M context, with agent-as-a-judge as new scaling axis
stochasticchasm · x · 2026-09-22
Commentary on Xiaomi MiMo's training ("classic mimo architecture, simple and strong") highlights two notable points:
- This may be the first model RL-trained at a 1M-token context (the author flags uncertainty)
- Heavy use of agent-as-a-judge at scale stands out — other Chinese labs are doing the same, effectively opening a new axis of compute scaling during RL that makes reward judging more robust.
A follow-up speculates that the "groupwise" setup may mean one agent per group, enabling contrastive grading, but this remains unconfirmed.
More from Research
- Diag2Diag: AI generates measurements that hardware sensors can't capture — AnneliesGamble · 2026-09-22
- RL training detail: re-prefilling instead of PipelineRL's cached KV, with batch size framed as a GPU-utilization lever — stochasticchasm · 2026-09-22
- RL infra detail: re-prefill over PipelineRL's KV cache reuse, batch size for GPU utilization — stochasticchasm · 2026-09-22
- New psychology paper uses social identity to explain false beliefs in AI-era information environments — steverathje2 · 2026-09-22
- Standard GRPO at 1M scale: why no critic models, and what counts as "behaviors"? — stochasticchasm · 2026-09-22
- Adam-to-Muon mid-training switch sparks debate, seemingly contradicting Moonlight paper — stochasticchasm · 2026-09-22