Xiaomi open-sources ~7K RL environments, but only 989 were used to train the 9B MiMo model
teortaxesTex · x · 2026-09-26
Xiaomi released 7,000 RL training environments for MiMo on HuggingFace, spanning code, cyber, general, music, and web dev. Insider Dorialexander clarifies: only a 989-env selection was used to RL-train a 9B distilled model (not the big MiMo); rewards aren't self-contained (need a judge setup or their grader service + VLM); and the real value is in the general/envs directory (with docker), which contains a rarely-seen mix of real/simulated documents.
Related event: Xiaomi Open-Sources MiMo-V2.6 RL Dataset, Tops HuggingFace Trending(5 posts)→
More from Models
- Theo slams OpenRouter for using Jev: a non-reasoning classifier that can't gauge task complexity — intellectronica · 2026-09-26
- VraserX: Gemini 4 Pro may lead in game design, but OpenAI still wins on science — VraserX · 2026-09-26
- "AI Safety Is Pseudoscience" Debate Hinges on OpenAI's Opaque Multi-Agent Training — basedjensen · 2026-09-26
- OpenAI docs add telephony support, letting AI agents dial and talk on phone calls — imjustnewatai · 2026-09-26
- Researchers surface spurious probes across models: Sonnet 5 recommends green tea in evals, oolong in production — jankulveit · 2026-09-26
- Burning 200M tokens for 30M useful ones: what maxing out models teaches you — gaganghotra_ · 2026-09-26