Paper: Sparse RL can't find what a model never samples; dense rewards break the ceiling
iScienceLuvr · x · 2026-08-27
A paper demystifying RL post-training (RLVR) of language models in a simplified setup yields three key results:
- Result 1: Sparse RL cannot find behaviors the model never samples — no reward signal helps if the behavior isn't sampled.
- Result 2: Dense rewards break that ceiling.
- Result 3: "Spurious rewards" are really a story about the prompt set, not the reward model.
The framing: post-training is best understood as redistributing probability mass inside the pretrained distribution. The suggested measurement is to track the probability the model assigns to the desired behavior, plus the entropy of its output distribution, throughout training.
Project page and code are open-sourced.
More from Research
- Scripps Research wins $19.5M NSF grant for AI-powered autonomous chemistry lab — CatAstro_Piyush · 2026-08-27
- Anatomy of Company Brains: 4 shared components across 9 projects — femke_plantinga · 2026-08-27
- 411k cut-outs from Britannica: 29M param model segments historical illustrations — vanstriendaniel · 2026-08-27
- van der Schaar Lab: What gets hidden when medicine is built around the average patient? — MihaelaVDS · 2026-08-27
- VGGT-SLAM++: Complete Visual SLAM System with Sim(3) Backend — rsasaki0109 · 2026-08-27
- Researcher kalomaze: papers leaning on 'pass@512 solves GSM8K' stop real analysis — kalomaze · 2026-08-27