MiMo-V2.6 streams RL run: DeepSWE score jumps from 58.41 to 72.57, nearing SOTA
teortaxesTex · x · 2026-09-23
nrehiew posted a paper thread on MiMo-V2.6, the newest open model and the latest to stream its RL training run. Unlike DeepSeek v4.1 Flash's paper, it focuses on data curation and RL experimental results. The RL runs have completed: the pro model improved from 58.41 to 72.57 on DeepSWE, approaching the leaderboard best of 74 held by Astra, Gemini 3.8 Flash, and Opus 5.
Related event: Xiaomi MiMo Finishes RL Training, DeepSWE Score Jumps to 72.57(2 posts)→
More from Models
- MachgenAI Offers Free Minimax H3 Turbo Generations for Accounts With $25+ Balance — TheMoonMidas · 2026-09-23
- Researcher: People Who Flex AI Sycophancy Screenshots Tend to Treat Humans Badly — repligate · 2026-09-23
- On Opus 5.5: The Corpus Is Full of Summaries of Summaries; Primary Experience Is Scarce — mimi10v3 · 2026-09-23
- Tester calls Opus 5.5's visual design the best of any model tested — repligate · 2026-09-23
- Opus 5.5 makes Claude 'finally back': better writing, low cost, high speed — nabeelqu · 2026-09-23
- repligate: Sycophantic AI users are often those who punish disagreement — repligate · 2026-09-23