CMU's TailRL optimizes reward upper tails, escaping suboptimal RL solutions

ceciletamura · x · 2026-10-05

A CMU team (Salakhutdinov, Bagnell, Zanette et al.) published Tail-Likelihood Reinforcement Learning (TailRL) on arXiv.

The problem: RL optimizes average reward, but for generative policies two strategies with the same mean can differ greatly in the chance of producing rare high-reward rollouts — the tail coverage that determines the benefit of extra sampling at train and inference time.

Method:

Across object localization, maze navigation, GUI grounding, and code optimization, TailRL leverages rare high-reward samples to avoid suboptimal solutions and benefits more from extra inference samples.

Original post →

More from Research

Research channel →