Xiaomi MiMo achieves streaming large-scale LLM RL training, including a 1T-parameter model
stanfordnlp · x · 2026-09-18
A member of Percy Liang's Stanford NLP group shared that XiaomiMiMo, following Marin's large-scale LLM pre-training effort on the 535B-A23B MoE model trained on 18T tokens, has achieved streaming large-scale LLM RL training including a 1T total-parameter model.\n\nThe linked Marin Tracker page shows the 535B run at roughly 30% (5.4T of 18T tokens), with a live log of engineering details: a failed checkpoint save once stopped the job, Python's automatic memory cleanup firing at different moments across machines causes recurring slow-step "spells", and fixes are still pending review and deployment.
More from Infra
- King Charles Meets OpenAI, Anthropic, DeepMind and Nvidia Execs on AI Safety — eyishazyer · 2026-09-18
- PlanetScale's TIN beats Postgres GIN full-text search: 212ms vs 288s at p99 — DanielLockyer · 2026-09-18
- NVIDIA shows 100x faster scikit-learn spectral clustering with cuML — NVIDIA Developer · 2026-09-18
- ChatGPT desktop app leaks context to the cloud by default — here's how to swap in Ollama — Technovangelist · 2026-09-18
- Reddit thread: what max-context KV reservations actually cost beyond concurrency — werunm · 2026-09-18
- Dev Patches vLLM to Run DiffusionGemma, Live Evals Show It Ties on Smarts but Loses on Speed to APIs — bodonoghue85 · 2026-09-18