TailRL optimizes reward tails instead of mean, cutting samples 256x on GUI grounding
burny_tech · x · 2026-09-07
Researchers propose Tail-Likelihood RL (TailRL), which targets a flaw in standard RL: optimizing average reward suppresses rare but exceptional rollouts and hurts Best-of-k scaling. TailRL instead maximizes the log-probability of exceeding a randomly chosen reward threshold, weighting rare high-reward samples more.
- Its gradient is a harmonic mixture of Best-of-k gradients across all k
- Requires only a simple advantage-function tweak, compatible with existing RL pipelines
- On GUI grounding it needs up to 256x fewer samples; code optimization Best-of-1024 sees a 7.7x speedup
- Evaluated on object localization, maze navigation, GUI grounding, and code optimization; code released as Zanette-Labs/TailRL
More from Research
- Computer-Use Models Are Still 'Low-Frequency, Highly Batched' — Minecraft May Stay Unsolved Until 2030 — mike64_t · 2026-09-07
- IBM's STAIR uses tables of contents for generative retrieval, hitting 82.6% Recall@1 — omarsar0 · 2026-09-07
- MinHash classic: clustering huge datasets with a KV store in five lines — moultano · 2026-09-07
- Safe RL With Stability Guarantees: Learning Without Ever Falling Down — tomssilver · 2026-09-07
- Principia benchmark: video models score 0.8 on VBench but under 0.42 on physics consistency — CSProfKGD · 2026-09-07
- 1-bit quantized embeddings cut vector index storage up to 60x with <1% quality loss — burkov · 2026-09-07