Achieving 0 Train-Infer Mismatch for MoE RL, Boosting Wordle Performance

PandaAshwinee · x · 2026-08-18

The team announced achieving zero numerical mismatch between training and inference for RL on large MoE models. This method significantly improved performance on the task of teaching Qwen3.6-35B-A3B to play Wordle. The code is open-source, and multiple ablation studies were conducted to verify the results.

Original post →

More from Infra

Infra channel →