Qwen MoE stability paper shows training-inference KL divergence rising over steps
heghbalz · x · 2026-07-24
- The post points to a Qwen MoE stability paper and highlights a plot of training-inference KL divergence.
- The author says they are trying to understand the intuition behind why the KL divergence between the trainer policy and inference grows over time, even though the experiments are synchronous and not due to async/off-policy effects.
- The attached chart shows KL divergence trending upward across gradient steps, with one run diverging sharply early.
More from Research
- AI Researcher: Recursive Self-Improvement Bottlenecked by Ecosystem Data — herbiebradley · 2026-07-24
- Paper shows bounded approximations for Fisher-Rao distance in parametric models — FrnkNlsn · 2026-07-24
- A “lol Grok” post surfaces the alignment debate over whether RL gives models real goals — Sauers_ · 2026-07-24
- Science spotlights 4D nucleome papers in a new issue featuring single-cell 3D genome work — jmuiuc · 2026-07-24
- CogSci 2026 best paper goes to “On convexity and efficiency in semantic systems” — gregd_nlp · 2026-07-24
- Proba predicts how fine-tuning data will change open-weight models before training — markjeffrey · 2026-07-24