RL Gains Concentrate on Easy Questions; 'Never Give Up' Resampling Tackles the Hard Ones

teortaxesTex · x · 2026-09-16

Researchers led by Nat Lambert find that RL gains on LLMs concentrate mostly on easy questions — a phenomenon they call the Matthew Effect for RL. Their fix is simple: when a GRPO group contains only wrong completions (zero gradient), resample with 0.9 probability until finding batches with nonzero gradient. Called "Never Give Up", the technique pairs with async RL to allocate more compute to harder problems, and it works. Paper and blog linked.

Related event: Research Reveals RL Matthew Effect; NGU Sampling Targets Hard Problems(4 posts)→

Original post →

More from Research

Research channel →