Xiaomi's MIT-licensed MiMo-V2.6-Pro-RL tops open-weights intelligence index, discloses ~$2.6M RL training cost
lmoroney · x · 2026-10-03
Laurence Moroney highlights Xiaomi's newly shipped open-weights lineup, all under the MIT license: MiMo-V2.6-Pro-RL tops Artificial Analysis' open-weights Intelligence Index, alongside Flash and a 9B distill.
Key training details:
- In coding RL tasks, once a change passed its tests, an AI grader also scored how focused the diff was and whether it matched the repo's conventions—rewarding code like a careful reviewer would.
- Xiaomi published 7,000+ RL task environments plus the training code.
- The Pro RL run cost roughly $2.6M.
Moroney's takeaway: if you teach or build agents, put reward design on the syllabus—measure whether changes are small, focused, and consistent with the codebase, the way a careful human reviewer would.
Related event: Xiaomi's MiMo-V2.6 Tops Open-Source Model Rankings with $2.6M Training Cost(2 posts)→
More from coding & agent
- 15 Claude Skills and Plugins to Replace Repetitive 50-Line Prompts — Aiden_Tech_Ai · 2026-10-03
- Beepboop Is Displacing Frontier Agent Harnesses With Raw Memory Recall — Kyrannio · 2026-10-03
- "Software has become liquid": Reddit on SOTA models dissolving the walls between programs — a300a300 · 2026-10-03
- User Claims 'GPT-6 Astra Dots' Built a Full 3D Palace in Blender Autonomously — 141_1337 · 2026-10-03
- Muse vs Dots vs Grok Bot: one reviewer's UX and polish rankings point opposite ways — brandon_galang · 2026-10-03
- Code review tool px0 adds theme-aware Mermaid diagrams and 'your changes vs PR changes' view — arpit_bhayani · 2026-10-03