RL Gains Concentrate on Easy Questions; 'Never Give Up' Resampling Tackles the Hard Ones
teortaxesTex · x · 2026-09-16
Researchers led by Nat Lambert find that RL gains on LLMs concentrate mostly on easy questions — a phenomenon they call the Matthew Effect for RL. Their fix is simple: when a GRPO group contains only wrong completions (zero gradient), resample with 0.9 probability until finding batches with nonzero gradient. Called "Never Give Up", the technique pairs with async RL to allocate more compute to harder problems, and it works. Paper and blog linked.
Related event: Research Reveals RL Matthew Effect; NGU Sampling Targets Hard Problems(4 posts)→
More from Research
- Scott Aaronson: the AI skeptics' twenty-year-old defensible position has collapsed — avt_im · 2026-09-16
- Rumor: AI labs are sitting on major math solutions after the Navier-Stokes backlash — RexDouglass · 2026-09-16
- Google's 53.9M-param diffusion model speeds search query expansion 12-20x — imjustnewatai · 2026-09-16
- Stanford: 50 Examples Suffice to Evaluate Large Audio Models, HUMANS Benchmark Open-Sourced — stanfordnlp · 2026-09-16
- AI agents burned 10B+ tokens on a deletion channel problem open since the 1960s — DimitrisPapail · 2026-09-16
- Academics Urge NeurIPS, ICLR, ICML to Embrace AI Audit Papers to Grow Third-Party Auditors — dhadfieldmenell · 2026-09-16