Xiaomi open-sources ~7K RL environments, but only 989 were used to train the 9B MiMo model

teortaxesTex · x · 2026-09-26

Xiaomi released 7,000 RL training environments for MiMo on HuggingFace, spanning code, cyber, general, music, and web dev. Insider Dorialexander clarifies: only a 989-env selection was used to RL-train a 9B distilled model (not the big MiMo); rewards aren't self-contained (need a judge setup or their grader service + VLM); and the real value is in the general/envs directory (with docker), which contains a rarely-seen mix of real/simulated documents.

Related event: Xiaomi Open-Sources MiMo-V2.6 RL Dataset, Tops HuggingFace Trending(5 posts)→

Original post →

More from Models

Models channel →