Tail-Likelihood RL paper targets upper-tail rewards that mean-based RL masks
burkov · x · 2026-09-07
RL for generative models typically maximizes expected reward, but policies with similar means can differ sharply in the probability of producing rare, high-value outputs — a gap that becomes decisive when sampling scales up during training or inference.
The paper introduces Tail-Likelihood Reinforcement Learning for continuous rewards: rather than optimizing a single mean, it treats exceeding each reward threshold as a binary success and maximizes the expected log-probability of success across thresholds drawn uniformly from the reward range. The resulting objective decomposes into a harmonic mixture of Best-of-k gradients, preserving upper-tail coverage while improving performance.
Related event: TailRL: Reinforcement Learning That Optimizes Tail Rewards(3 posts)→
More from Research
- Thousands of AI drug discovery firms, only a handful of end-to-end automated labs — robleclerc · 2026-09-08
- Insilico's AI drug Rentosertib shows 3-4 year biological age reversal in 12 weeks — NinaDSchick · 2026-09-08
- Speridlabs releases ENEAS, a text-promptable tracking method claiming gains over SAM3 — joecole · 2026-09-08
- UCL's Large Discovery Models hit SOTA, 2.4x LLM reflection on experiment search — jiqizhixin · 2026-09-08
- First IAB workshop on agent behavior draws 260 submissions, seeks reviewers — mdredze · 2026-09-08
- Dr. Claw: Open-Source AI Scientist Workspace Wrapping Coding Agents for Auditable Research — Dingjie Song · 2026-09-08