Xiaomi MiMo may be first RL at 1M context, with agent-as-a-judge as new scaling axis

stochasticchasm · x · 2026-09-22

Commentary on Xiaomi MiMo's training ("classic mimo architecture, simple and strong") highlights two notable points:

A follow-up speculates that the "groupwise" setup may mean one agent per group, enabling contrastive grading, but this remains unconfirmed.

Related event: Xiaomi MiMo: Possibly First 1M-Context RL Training, Big Gains in Just 30 Steps(4 posts)→

Original post →

More from Research

Research channel →