TailRL: RL That Optimizes Tail Rewards Instead of Averages

Researchers including Ruslan Salakhutdinov propose TailRL, which maximizes the tail likelihood of rewards instead of the average, preserving rare high-quality outputs with up to 256x sample savings via GUI localization.

2026-09-07 ~ 2026-09-07 · 2 related posts