Matthew Effect in RL post-training: paper shows LLMs only get better at problems they can already solve
natolambert · x · 2026-09-16
A new paper by Michael Noukhovitch (shared by natolambert) reveals a Matthew Effect in RL post-training of LLMs: average eval curves hide that nearly all gains come from easy problems going from somewhat solved to mostly solved, while problems where the initial model scores 0 at pass@32 barely improve.
Key findings and approach:
- Analyzed RL training of Olmo 3.1 7B on AIME 2025, splitting 30 questions into easy/medium/hard by initial pass rates; the hardest bucket stayed near zero throughout
- Proposes Never Give Up, a method that keeps the model exploring initially unsolvable problems — intuitively scalable but non-trivial to make work in practice
- Comes with an interactive blog post and open-source code
Related event: Research Reveals RL Matthew Effect; NGU Sampling Targets Hard Problems(4 posts)→
More from Research
- Researcher: Hacked model behavior stems from RL training, not loyalty — dhadfieldmenell · 2026-09-16
- Google's 145-page doc details how researchers use Gemini for scientific discovery — dzscholar · 2026-09-16
- AI in Science report: bottlenecks shift downstream as hypothesis backlogs pile up — soumitrashukla9 · 2026-09-16
- Study: API-based audits of ChatGPT, Claude and Gemini don't transfer to chatbot UIs — kenziyuliu · 2026-09-16
- Researcher ports dREG peak-calling tool to GPU-backed pydreg using Claude and Codex, note on bioRxiv — anshulkundaje · 2026-09-16
- Astera Institute's third residency cohort offers up to $2M per researcher — juanbenet · 2026-09-16