TailRL: New RL Objective Maximizes Upper-Tail Reward Coverage Instead of Just the Mean
burny_tech · x · 2026-09-07
Ruslan Salakhutdinov and collaborators released Tail-Likelihood Reinforcement Learning (TailRL), a new paper asking whether RL is optimizing the right objective.
- Motivation: Standard RL optimizes mean reward, which often collapses the policy distribution to a spike and kills probability mass on rare but high-reward rollouts — exactly the samples that matter as inference-time sampling scales up.
- Method: TailRL turns a continuous reward into a family of binary success events (exceeding a randomly chosen threshold) and maximizes the expected log-probability of exceeding it, jointly boosting mean reward and coverage of high-reward outputs. Its gradient weights rare high-reward rollouts more and can be interpreted as a mixture of Best-of-k gradients.
- Drop-in: Only a simple modification to the advantage function, compatible with existing RL pipelines.
- Results: Across object localization, maze navigation, GUI grounding, and code optimization, TailRL exploits rare high-reward samples and scales better with increased sampling.
Paper: arXiv 2609.02987, with authors including Andrea Zanette, Aarti Singh, and J. Andrew Bagnell.
Related event: TailRL: RL That Optimizes Tail Rewards Instead of Averages(2 posts)→
More from Research
- MobileWorld benchmark: 201-task eval of phone GUI agents, best combo only ~52% — East-Muffin-6472 · 2026-09-07
- saprmarks: Supervised Activation Probes Aren't Mechanistic Interpretability Either — saprmarks · 2026-09-07
- Martian routes across 44 LLMs, cutting errors 46% vs the best single model — HowDevelop · 2026-09-07
- Hameroff revives quantum consciousness theory: aromatic amino acids form a 'Quantum Underground' — JosephJacks_ · 2026-09-07
- ShallowStream cuts streaming video understanding latency with shallow-layer index then deep answering — Jitai Hao · 2026-09-07
- Stanford's Chris Potts revives 8-year-old "Deep RL Doesn't Work Yet" to counter AI skeptics — ChrisGPotts · 2026-09-07