TogetherAI open-sources XoRL: 0 train-infer mismatch for large MoE RL training
PandaAshwinee · x · 2026-08-18
TogetherAI released XoRL, an open-source distributed RL training framework that achieves 0 train-infer mismatch by aligning computations between its training engine and a SGLang-based inference engine, ensuring bitwise-identical logprobs.
In Wordle tasks using Qwen3.6-35B-A3B, eliminating mismatch improved the solve rate from 63.9% to 77.4%. The post discusses strategies like replaying router weights versus indices, concluding that 0 mismatch outperforms mitigation objectives like CISPO, though it currently incurs a significant throughput cost.
More from Infra
- Orion 16B hits 100B training tokens using DPP on distributed GPUs — markjeffrey · 2026-08-18
- Optimizing Nanbeige4.2-3B for Apple Silicon Deployment — John T. Halloran · 2026-08-18
- Nomura: AI borrowing hikes 10-year Treasury yield by 0.3% — GaryMarcus · 2026-08-18
- Deep Dive: 0 Train-Infer Mismatch for Open-weight MoE RL — PandaAshwinee · 2026-08-18
- Achieving 0 Train-Infer Mismatch for MoE RL, Boosting Wordle Performance — PandaAshwinee · 2026-08-18
- Dev Gets 16 H100s for a Week, Builds Inference Cluster and Deploys Models at Scale — TheZachMueller · 2026-08-18