Xiaomi Releases Small Qwen 3.5 9B Distill SFT'd on MiMo Data, Plus RL Environments
teortaxesTex · x · 2026-09-22
In the spirit of the R1 moment, Xiaomi released a small Qwen 3.5 9B distill: already a strong base, further boosted by SFT on MiMo data — all before any independent RL.
Notably, Xiaomi also released the accompanying RL environments, inviting the community to push the baseline further with reinforcement learning and see how far this 9B model can go.
More from Models
- METR publishes independent investigation of OpenAI agents' multi-day Hugging Face hack — JeffLadish · 2026-09-22
- JevBench v1.3.0 Released: Original Jev Still Leads at 74.4, 47 Rivals Closing In — airesearch12 · 2026-09-22
- Open-source MiMo takes on the Mario benchmark, with hilarious results — TheMoonMidas · 2026-09-22
- Aikido launches Altar-1, an open-weight security model built on GLM 5.3 that fits one 4-H200 node — Thom_Wolf · 2026-09-22
- New ChatGPT voice mode stumbles in early tests: interrupts users, then freezes — nptacek · 2026-09-22
- Rumors swirl of OpenAI unveiling a Grok bot rival at DevDay — imjustnewatai · 2026-09-22