Alibaba proposes I-SDPO to fix degenerate gradient problem in GRPO
heghbalz · x · 2026-08-17
Alibaba's latest paper introduces the I-SDPO framework to tackle the degenerate gradient problem in GRPO. When prompts are tough and every response fails, the policy gradient flatlines. While self-distillation helps, it can lock the model into a rigid style. I-SDPO makes distillation adaptive at the instance level: it uses privileged self-distillation only if an entire rollout group fails, flipping back to standard GRPO the moment a single attempt succeeds.
More from Research
- Retriever: A Framework for Asynchronous, Closed-Loop Robot Agents — ZeYanjie · 2026-08-24
- Converting GMMs ↔ PEFs for fast KLD approximation — FrnkNlsn · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24