Dense Rewards Break RL Post-Training Ceiling, Sparse Rewards Fail
joecole · x · 2026-08-28
This post discusses Reinforcement Learning (RL) in LLM post-training. It argues that sparse rewards (0 or 1) are ineffective and cannot find behaviors the model never samples. In contrast, continuous dense rewards (0 to 1) can break the ceiling set by the pre-training distribution. Post-training is framed as redistributing probability mass within that distribution, suggesting that tracking the probability of desired behavior and output entropy is crucial.
Related event: Dense Rewards, Not Sparse Ones, Break Post-Training Ceilings(2 posts)→
More from Research
- ProgramBench: Factory's benchmark makes agents reproduce real software from scratch — shaunmmaguire · 2026-08-28
- Sai tops OSWorld 2.0 benchmark, beats GPT-5.6 and Opus 5 at lower cost — taoyds · 2026-08-28
- Factory releases ProgramBench: A benchmark for reproducing real software from scratch — matanSF · 2026-08-28
- Ablation discussion: query head count barely matters early, sparse routing must be learned — stochasticchasm · 2026-08-28
- 404 Team to Walkthrough Titan Model and World Models — markjeffrey · 2026-08-28
- NVIDIA Details QAD Pipeline for Optimizing Nemotron Model — PyTorch · 2026-08-28