Bellman Policy Optimization: critic-free RLVR method for LLM reasoning

apodex · hf · 2026-09-23

The paper introduces Bellman Policy Optimization (BPO), a critic-free RLVR method derived from Policy Mirror Descent. Using Bellman equations, BPO reformulates PMD as a trajectory-level objective for autoregressive generation with terminal rewards, avoiding intermediate state-value estimation, and proves equivalence of optimal solutions. The practical loss uses a mismatch-correction weight based on a smoothed ratio of complementary token probabilities. Experiments on math reasoning benchmarks show its effectiveness.

Original post →

More from Research

Research channel →