AI2's NGU sampling fixes RL for LLMs that only improves easy tasks
allenai · hf · 2026-09-16
Allen AI published research showing that RL for LLMs disproportionately improves easy tasks while hard problems barely benefit. They propose NGU (Never Give Up), an adaptive sampling method that detects persistently unsolved problems and reallocates compute from already-mastered easy ones to these hard samples, boosting overall performance. The work highlights sample allocation as an underrated lever in RLVR and reasoning-model training.
More from Models
- Grok 4.6 matches Opus 5 on 3D reconstruction at 1/5 the cost, testers claim — Lianhuiq · 2026-09-16
- Exclusive: Open Chinese models close gap with Silicon Valley's frontier AI — alexvoica · 2026-09-16
- Lithos: stop treating AI benchmarks as proof, define your own metrics — JiaZhihao · 2026-09-16
- Claude Opus 5 'Completely Useless Today': Developer Slams Model for Year-Old-Level Mistakes — jasonkneen · 2026-09-16
- Rumor: SSI cracked test-time training with a new scaling law; combined with multi-agent could yield AGI — toptickcrypto · 2026-09-16
- Ex-OpenAI researcher launches TypeSafe's Jev: System One models 100x faster, no hallucination — rohanpaul_ai · 2026-09-16