Paper: Sparse RL can't find what a model never samples; dense rewards break the ceiling

iScienceLuvr · x · 2026-08-27

A paper demystifying RL post-training (RLVR) of language models in a simplified setup yields three key results:

The framing: post-training is best understood as redistributing probability mass inside the pretrained distribution. The suggested measurement is to track the probability the model assigns to the desired behavior, plus the entropy of its output distribution, throughout training.

Project page and code are open-sourced.

Original post →

More from Research

Research channel →