Bellman Policy Optimization: critic-free RLVR method for LLM reasoning
apodex · hf · 2026-09-23
The paper introduces Bellman Policy Optimization (BPO), a critic-free RLVR method derived from Policy Mirror Descent. Using Bellman equations, BPO reformulates PMD as a trajectory-level objective for autoregressive generation with terminal rewards, avoiding intermediate state-value estimation, and proves equivalence of optimal solutions. The practical loss uses a mismatch-correction weight based on a smoothed ratio of complementary token probabilities. Experiments on math reasoning benchmarks show its effectiveness.
More from Research
- Shane Legg's visionary 2008 PhD dissertation 'Machine Super Intelligence' resurfaces online — burkov · 2026-09-23
- Naproche project site: formal math in natural, mathematician-readable language — zetalyrae · 2026-09-23
- Naproche: a proof assistant that reads math proofs written in controlled natural language — zetalyrae · 2026-09-23
- Nanjing University used AI to design proteins 4x stronger than any natural protein — MikePFrank · 2026-09-23
- Open-source GUI agents top out at 8% task success on composite cross-device tasks — maier_ak · 2026-09-23
- JarvisGUI benchmark tests GUI agents across Android, Windows and Ubuntu in one workflow — maier_ak · 2026-09-23