Research Reveals RL Matthew Effect; NGU Sampling Targets Hard Problems

New research from Allen AI reveals a 'Matthew effect' in RL training of LLMs, where gains concentrate on easy problems. The team proposes adaptive 'Never Give Up' sampling to reallocate compute toward hard problems.

2026-09-16 ~ 2026-09-16 · 4 related posts