TogetherAI Open-Sources XoRL for Zero Train-Inference Mismatch in MoE RL
TogetherAI open-sourced XoRL, a distributed RL framework that achieves zero train-inference mismatch for large MoE models by tightly aligning computation at the kernel level. It significantly improved Wordle task performance with Qwen3.6-35B-A3B, with River API outperforming Tinker.
2026-08-18 ~ 2026-08-18 · 4 related posts
- TogetherAI open-sources XoRL: 0 train-infer mismatch for large MoE RL training — PandaAshwinee · 2026-08-18
- Achieving 0 Train-Infer Mismatch for MoE RL, Boosting Wordle Performance — PandaAshwinee · 2026-08-18
- Open Source Method Eliminates Train-Inference Mismatch in MoE RL, Boosting Performance — PandaAshwinee · 2026-08-18
- River API outperforms Tinker in RL, enabling zero train-infer mismatch for MoEs — marcbhargava · 2026-08-18