AI2's NGU sampling fixes RL for LLMs that only improves easy tasks

allenai · hf · 2026-09-16

Allen AI published research showing that RL for LLMs disproportionately improves easy tasks while hard problems barely benefit. They propose NGU (Never Give Up), an adaptive sampling method that detects persistently unsolved problems and reallocates compute from already-mastered easy ones to these hard samples, boosting overall performance. The work highlights sample allocation as an underrated lever in RLVR and reasoning-model training.

Original post →

More from Models

Models channel →