Qwen MoE stability paper shows training-inference KL divergence rising over steps
heghbalz · x · 2026-07-24
- The post points to a Qwen MoE stability paper and highlights a plot of training-inference KL divergence.
- The author says they are trying to understand the intuition behind why the KL divergence between the trainer policy and inference grows over time, even though the experiments are synchronous and not due to async/off-policy effects.
- The attached chart shows KL divergence trending upward across gradient steps, with one run diverging sharply early.
More from Research
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11