Soft Data Duplication Blows Out Gradients: Root Cause of GRPO Training Collapse
ivan_bezdomny · x · 2026-08-02
The author diagnosed the root cause of the previous GRPO training collapse in Gemma. GRPO generates multiple outputs per input to compute gradients. When consecutive batches contain many of the same rare tokens (e.g., a specific bill name), it creates gradients in the exact same direction. This double-step backprop, amplified by momentum, blows out the model weights irrecoverably.
Related event: Developer Reveals LLM Fine-Tuning Fragility and Training Crashes(3 posts)→
More from Research
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24
- Study: Agents read instructions/notes 60.5% of the time, rarely touch API docs — dair_ai · 2026-08-24