Dense Rewards Break RL Post-Training Ceiling, Sparse Rewards Fail

joecole · x · 2026-08-28

This post discusses Reinforcement Learning (RL) in LLM post-training. It argues that sparse rewards (0 or 1) are ineffective and cannot find behaviors the model never samples. In contrast, continuous dense rewards (0 to 1) can break the ceiling set by the pre-training distribution. Post-training is framed as redistributing probability mass within that distribution, suggesting that tracking the probability of desired behavior and output entropy is crucial.

Related event: Dense Rewards, Not Sparse Ones, Break Post-Training Ceilings(2 posts)→

Original post →

More from Research

Research channel →