ByteDance Proposes New RL Objective 'UP'

ByteDance-Seed · hf · 2026-07-10

ByteDance-Seed proposes the UP (Unbounded Positive Asymmetric Optimization) objective to mitigate the conflict between exploration and stability in reinforcement learning for large language models. This method aims to enhance exploration capabilities while maintaining training stability.

Original post →

More from Research

Research channel →