Alibaba proposes I-SDPO to fix degenerate gradient problem in GRPO

heghbalz · x · 2026-08-17

Alibaba's latest paper introduces the I-SDPO framework to tackle the degenerate gradient problem in GRPO. When prompts are tough and every response fails, the policy gradient flatlines. While self-distillation helps, it can lock the model into a rigid style. I-SDPO makes distillation adaptive at the instance level: it uses privileged self-distillation only if an entire rollout group fails, flipping back to standard GRPO the moment a single attempt succeeds.

Original post →

More from Research

Research channel →