ByteDance Proposes New RL Objective 'UP'
ByteDance-Seed · hf · 2026-07-10
ByteDance-Seed proposes the UP (Unbounded Positive Asymmetric Optimization) objective to mitigate the conflict between exploration and stability in reinforcement learning for large language models. This method aims to enhance exploration capabilities while maintaining training stability.
More from Research
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- PNAS special issue examines copyright, governance, and AI in the legal system — chrmanning · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22