TailRL optimizes reward tails instead of mean, cutting samples 256x on GUI grounding

burny_tech · x · 2026-09-07

Researchers propose Tail-Likelihood RL (TailRL), which targets a flaw in standard RL: optimizing average reward suppresses rare but exceptional rollouts and hurts Best-of-k scaling. TailRL instead maximizes the log-probability of exceeding a randomly chosen reward threshold, weighting rare high-reward samples more.

Original post →

More from Research

Research channel →