Xiaomi livestreams MiMo-v2.6 RL training run, hits ~60% of deepswe in one day
zainhas · x · 2026-09-17
Xiaomi is publicly livestreaming the full reinforcement learning run for its MiMo-v2.6 and 2.6 flash models. Per the poster, the run started just one day ago and has already reached roughly 60% progress on the deepswe benchmark — an unusually transparent move for a major model training effort.
Related event: Xiaomi live-streams MiMo-V2.6 RL training with public dashboard(5 posts)→
More from Models
- Burkov predicts looping recurrent 7B transformers will return and get good at coding — burkov · 2026-09-17
- OpenAI's big 'ship week' reportedly postponed, GPT-6 Sol timing now unclear — testingcatalog · 2026-09-17
- Mozilla's 91-page report: open-weight AI now only ~4 months behind the frontier — rohanpaul_ai · 2026-09-17
- Stealth model leak speculated to be Mistral: no output-token billing for reasoning, answers China questions — zainhas · 2026-09-17
- Unreleased Astra-family model reportedly developed a new persona banner during RL training — inductionheads · 2026-09-17
- Researcher Despairs as Gemini Cites 'Emergent Mind' for Made-up AUROC Baselines — anshulkundaje · 2026-09-17