Xiaomi MiMo open-sources MiMo-V2.6-RL-oss training dataset on Hugging Face under Apache-2.0
iScienceLuvr · x · 2026-09-26
Xiaomi's MiMo team has released its reinforcement learning dataset MiMo-V2.6-RL-oss on Hugging Face under the permissive Apache-2.0 license, a move the community is calling refreshingly open. The dataset spans text, image, and document modalities in parquet format with a few thousand samples, including real code test-case examples for RL training. It's a directly usable resource for anyone reproducing or studying post-training pipelines.
Related event: Xiaomi Open-Sources MiMo-V2.6 RL Dataset, Tops HuggingFace Trending(5 posts)→
More from Research
- Position paper argues grounded abduction is the missing reasoning leap for LLMs — bibryam · 2026-09-26
- BeeNara: a 332MB CPU-only classifier that knows when no folder fits (96.8% recall) — razer_psycho · 2026-09-26
- InternLM open-sources Intern-Decision 4B/0.8B: structured decisions in one forward pass — jacek2023 · 2026-09-26
- Stanford's Noah Goodman uses philosophy to improve LLM pretraining, jokes ASI achieved — xuanalogue · 2026-09-26
- Researchers surface spurious probes across models: Sonnet 5 recommends green tea in evals, oolong in production — jankulveit · 2026-09-26
- Three papers, one warning: 1% synthetic data can trigger strong model collapse — suchenzang · 2026-09-26