Xiaomi open-sources the RL environment and code used to train MiMo v2.6 pro/flash
zainhas · x · 2026-09-26
- Xiaomi open-sourced on Hugging Face the environment and code used to RL-train MiMo v2.6 pro/flash, as the dataset XiaomiMiMo/MiMo-V2.6-RL-oss under Apache-2.0.
- The release includes document, image and text modality RL training data, letting the community reproduce or learn from its reinforcement learning training pipeline.
Related event: Xiaomi Open-Sources 7,000+ RL Training Environments for MiMo(3 posts)→
More from Research
- Alibaba's Ovis-Embedding maps text, images, video and audio into one space, SOTA on MMEB-v3 — solyarisoftware · 2026-09-27
- Pretraining is just RL with single-token rollouts and full-information feedback — burny_tech · 2026-09-27
- Language is a lossy compression: even the best models train on a thin residue of reality — yunta_tsai · 2026-09-27
- Maybe the nativist-empiricist controversy boils down to two meanings of the polysemous word 'learn' — abenitezburraco · 2026-09-27
- Valid JSON isn't a valid decision: LLM output consistency measured as low as 14.4% — tenkei_01 · 2026-09-27
- A Historical Introduction to Self-Play in Reinforcement Learning, Explained — cephaloform · 2026-09-27