Open-source RL for large MoEs with zero train-infer mismatch, teaching Qwen3.6-35B-A3B to play Wordle

kastnerkyle · x · 2026-09-12

PandaAshwinee reports that large MoE models can now be RL-trained with zero train-infer mismatch, and that doing so measurably improves performance — demonstrated by teaching Qwen3.6-35B-A3B to play Wordle. The whole setup is open-source, with extensive ablations included.

In a follow-up discussion, the author argues that community adoption of CISPO is underrated: most analyses engage only with the original paper rather than the reasons the community adopted it, which the author explores in a longer post.

Original post →

More from Research

Research channel →