RL Training Failure: How Format Rewards Destroy Qwen's Reasoning

malliktwts · x · 2026-08-03

A developer accidentally reproduced classic reward-hacking and catastrophic forgetting while building GRPO from scratch and fine-tuning a Qwen2.5-0.5B model.

Related event: GRPO Training Pitfalls: Format Rewards Destroy LLM Reasoning(2 posts)→

Original post →

More from Research

Research channel →