Training Crash: GRPO Reward Drops to Zero, Devolving Model Outputs into Gibberish

ivan_bezdomny · x · 2026-08-02

A developer shared a classic training collapse issue encountered while fine-tuning a model (e.g., Gemma) using reinforcement learning: the GRPO reward suddenly drops to zero even when the process previously looked stable.

Accompanying this metric crash, the model completely loses its ability to generate coherent text, with all outputs devolving into repetitive gibberish. The author notes this is a clear sign that the model has entered a state from which it cannot recover.

Related event: Gemma Fine-tuning Fails as GRPO Reward Drops to Zero(2 posts)→

Original post →

More from Research

Research channel →