Xiaomi MiMo Finishes RL Training, DeepSWE Score Jumps to 72.57

Xiaomi's MiMo completed reinforcement learning training, boosting its DeepSWE score from 58.41 to 72.57—near the 74 record—while MiMo-V2.6 emerged as the latest open-source model to publicly document its RL training.

2026-09-21 ~ 2026-09-23 · 2 related posts

Full story(3 episodes)→