Xiaomi MiMo achieves streaming large-scale LLM RL training, including a 1T-parameter model

stanfordnlp · x · 2026-09-18

A member of Percy Liang's Stanford NLP group shared that XiaomiMiMo, following Marin's large-scale LLM pre-training effort on the 535B-A23B MoE model trained on 18T tokens, has achieved streaming large-scale LLM RL training including a 1T total-parameter model.\n\nThe linked Marin Tracker page shows the 535B run at roughly 30% (5.4T of 18T tokens), with a live log of engineering details: a failed checkpoint save once stopped the job, Python's automatic memory cleanup firing at different moments across machines causes recurring slow-step "spells", and fixes are still pending review and deployment.

Original post →

More from Infra

Infra channel →