Qwen MoE stability paper shows training-inference KL divergence rising over steps
heghbalz · x · 2026-07-24
- The post points to a Qwen MoE stability paper and highlights a plot of training-inference KL divergence.
- The author says they are trying to understand the intuition behind why the KL divergence between the trainer policy and inference grows over time, even though the experiments are synchronous and not due to async/off-policy effects.
- The attached chart shows KL divergence trending upward across gradient steps, with one run diverging sharply early.
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11