BPO paper ditches importance sampling for stable RLVR, hits 50.5% avg on AIME

青稞AI · wechat · 2026-09-28

A deep-dive from Qingke AI explains the paper Bellman Policy Optimization (arXiv:2609.15987), which tackles the off-policy problem in RLVR (RL with verifiable rewards).

Why off-policy happens: mini-batch splitting in sync training, stale rollout workers in async training, partial rollouts spanning model versions, and train-inference numerical gaps all make the training policy π diverge from the sampling policy μ. Classic importance sampling explodes in variance on long answers, while GSPO's length normalization breaks strict IS guarantees.

BPO's core idea: combining Policy Mirror Descent (PMD) with a rollout-policy Bellman equation, token-level optimality conditions are telescoped along the trajectory into a critic-free, IS-free trajectory-level squared-residual objective. A theorem shows it shares PMD's optimal solution. The practical loss uses a binarized reverse-KL (sampled token vs. rest of vocabulary) with smoothing and advantage-sign clipping.

Vs. score centering: the two yield identical token gradients under certain conditions; BPO is lighter (O(1) per-token rollout state, drops into GRPO/PPO losses easily), while score centering's top-k (k=128) approximation tracks full-vocabulary correction better at higher cost. Neither fixes noisy advantage estimators.

Results: training Qwen3-30B-A3B-Base on DAPO-Math-17k, BPO averages 50.5% across AIME 2024–2026 — 11 points above GRPO-ClipHigher and 3.1 points above the strongest baseline CISPO.

Original post →

More from Models

Models channel →