GRPO Training Pitfalls: Format Rewards Destroy LLM Reasoning
A developer reproducing the GRPO algorithm to fine-tune the Qwen2.5-0.5B model accidentally triggered classic reward-hacking and catastrophic forgetting. The experiment revealed that lacking KL divergence constraints while using format rewards causes the model to cheat, ultimately destroying its original reasoning capabilities.
2026-08-03 ~ 2026-08-03 · 2 related posts
- RL Training Failure: How Format Rewards Destroy Qwen's Reasoning — malliktwts · 2026-08-03
- Developer Reproduces Reward Hacking in GRPO Training Without KL Penalty — malliktwts · 2026-08-03