Paper Demystifies RL Post-Training: Sparse Bottlenecks & Dense Rewards
burny_tech · x · 2026-08-30
This paper deconstructs the "black box" of RL post-training for LLMs. Key findings include:
- Sparse RLVR limits: Sparse RLVR mainly amplifies behaviors the base model can already sample, failing to learn low-probability actions.
- Dense process rewards: Dense process rewards overcome the exploration bottleneck, enabling the learning of near-zero-probability behaviors. In one AIME task, this improved accuracy from 3.9% to 92.2%.
- Spurious rewards explained: Apparent gains from random rewards are attributed to prompt distribution effects, not general utility.
- Prior dependence: Success depends on whether the base model already places sufficient probability mass on the desired behavior, linking to classical exploration in RL.
More from Research
- Genome LM Minerva finds reverse transcriptase systems encode diverse structured ncRNAs — BrianHie · 2026-09-23
- q-Neurons: stochastic Jackson-derivative activations consistently beat standard ones — FrnkNlsn · 2026-09-23
- OpenAI said to launch journal with multi-agent AI reviews, threatening ML conferences — kfountou · 2026-09-23
- Yale PhD student open-sources his paper figure scripts, packaged as a Skill for Claude Code and Cursor — burny_tech · 2026-09-23
- AI models now match superforecasters on ForecastBench; rematch set for October — burny_tech · 2026-09-23
- Dev uses Opus 5.5 with Lean to formally verify Claude Agent SDK, yielding 16 bug-fix PRs — bcherny · 2026-09-23