Xiaomi Releases Small Qwen 3.5 9B Distill SFT'd on MiMo Data, Plus RL Environments

teortaxesTex · x · 2026-09-22

In the spirit of the R1 moment, Xiaomi released a small Qwen 3.5 9B distill: already a strong base, further boosted by SFT on MiMo data — all before any independent RL.

Notably, Xiaomi also released the accompanying RL environments, inviting the community to push the baseline further with reinforcement learning and see how far this 9B model can go.

Original post →

More from Models

Models channel →