Open-source RL for large MoEs with zero train-infer mismatch, teaching Qwen3.6-35B-A3B to play Wordle
kastnerkyle · x · 2026-09-12
PandaAshwinee reports that large MoE models can now be RL-trained with zero train-infer mismatch, and that doing so measurably improves performance — demonstrated by teaching Qwen3.6-35B-A3B to play Wordle. The whole setup is open-source, with extensive ablations included.
In a follow-up discussion, the author argues that community adoption of CISPO is underrated: most analyses engage only with the original paper rather than the reasons the community adopted it, which the author explores in a longer post.
More from Research
- If AI can produce correct proofs cheaply, what still matters? A mathematician's take — RexDouglass · 2026-09-12
- Would a yes/no oracle for theories be useful? Yes — an answer alone collapses the search space — basedjensen · 2026-09-12
- Meta paper: adversarial persuasion flips 62-91% of LLM judge verdicts, 70% of flips drift from ground truth — rohanpaul_ai · 2026-09-12
- Navier-Stokes AI proof took 10,000 agents, 88 hours and 130B tokens — not superintelligence — Healthy_Outcome7897 · 2026-09-12
- VIGA agent rebuilds images into editable Blender scenes via multimodal inverse-graphics loop — Michael_J_Black · 2026-09-12
- How Can LLM RL Work Despite Information-Theoretic Inefficiency? A Deep Dive — nrehiew_ · 2026-09-12