Research Unravels GRPO Training Collapse Cause and Fix

VectorInst · x · 2026-07-08

Xiaoxiao Li's team investigated why the GRPO algorithm, widely used in reasoning model training pipelines, repeatedly crashes: models gradually lose confidence in their own outputs, leading to gradient instability. Their proposed fix delivers performance improvements of up to 45% on key benchmarks.

Related event: Research Unravels GRPO Training Collapse; ICML Highlights AI Scaling(2 posts)→

Original post →

More from Research

Research channel →