TogetherAI open-sources XoRL: 0 train-infer mismatch for large MoE RL training

PandaAshwinee · x · 2026-08-18

TogetherAI released XoRL, an open-source distributed RL training framework that achieves 0 train-infer mismatch by aligning computations between its training engine and a SGLang-based inference engine, ensuring bitwise-identical logprobs.

In Wordle tasks using Qwen3.6-35B-A3B, eliminating mismatch improved the solve rate from 63.9% to 77.4%. The post discusses strategies like replaying router weights versus indices, concluding that 0 mismatch outperforms mitigation objectives like CISPO, though it currently incurs a significant throughput cost.

Original post →

More from Infra

Infra channel →