CMU's TailRL optimizes reward upper tails, escaping suboptimal RL solutions
ceciletamura · x · 2026-10-05
A CMU team (Salakhutdinov, Bagnell, Zanette et al.) published Tail-Likelihood Reinforcement Learning (TailRL) on arXiv.
The problem: RL optimizes average reward, but for generative policies two strategies with the same mean can differ greatly in the chance of producing rare high-reward rollouts — the tail coverage that determines the benefit of extra sampling at train and inference time.
Method:
- Convert continuous reward into a family of binary success events: for each threshold, how likely is the policy to exceed it
- TailRL maximizes the log-probability of exceeding a randomly chosen threshold; its gradient upweights rare high-reward rollouts and can be read as a mixture of Best-of-k gradients
- Only a simple change to the advantage function, compatible with existing RL pipelines
Across object localization, maze navigation, GUI grounding, and code optimization, TailRL leverages rare high-reward samples to avoid suboptimal solutions and benefits more from extra inference samples.
More from Research
- Dream4ACT Unifies Video-Action Modeling Across Robot Embodiments, Hitting 89% on RoboTwin 2.0 — Xiangyu Zhu · 2026-10-05
- SMI Brings Understanding-Driven Spatial Memory Management to Long-Video World Models — Ying Yang · 2026-10-05
- LVMT Sets New SOTA in Long-Term Video Segmentation While Running 10X Faster — tue-mps · 2026-10-05
- Schmidhuber: Google's 2017 Transformer Builds on His 1991 Linear Attention Work — SchmidhuberAI · 2026-10-05
- Near-identical image scores, huge gaps: AI denoising must serve science, not looks — bravo_abad · 2026-10-05
- SDECast: Neural SDEs Bring Continuous-Time Probabilistic Weather Forecasts Out to 5 Days — canaesseth · 2026-10-05