MiMo-V2.6 streams RL run: DeepSWE score jumps from 58.41 to 72.57, nearing SOTA

teortaxesTex · x · 2026-09-23

nrehiew posted a paper thread on MiMo-V2.6, the newest open model and the latest to stream its RL training run. Unlike DeepSeek v4.1 Flash's paper, it focuses on data curation and RL experimental results. The RL runs have completed: the pro model improved from 58.41 to 72.57 on DeepSWE, approaching the leaderboard best of 74 held by Astra, Gemini 3.8 Flash, and Opus 5.

Related event: Xiaomi MiMo Finishes RL Training, DeepSWE Score Jumps to 72.57(2 posts)→

Original post →

More from Models

Models channel →