BPO paper ditches importance sampling for stable RLVR, hits 50.5% avg on AIME
青稞AI · wechat · 2026-09-28
A deep-dive from Qingke AI explains the paper Bellman Policy Optimization (arXiv:2609.15987), which tackles the off-policy problem in RLVR (RL with verifiable rewards).
Why off-policy happens: mini-batch splitting in sync training, stale rollout workers in async training, partial rollouts spanning model versions, and train-inference numerical gaps all make the training policy π diverge from the sampling policy μ. Classic importance sampling explodes in variance on long answers, while GSPO's length normalization breaks strict IS guarantees.
BPO's core idea: combining Policy Mirror Descent (PMD) with a rollout-policy Bellman equation, token-level optimality conditions are telescoped along the trajectory into a critic-free, IS-free trajectory-level squared-residual objective. A theorem shows it shares PMD's optimal solution. The practical loss uses a binarized reverse-KL (sampled token vs. rest of vocabulary) with smoothing and advantage-sign clipping.
Vs. score centering: the two yield identical token gradients under certain conditions; BPO is lighter (O(1) per-token rollout state, drops into GRPO/PPO losses easily), while score centering's top-k (k=128) approximation tracks full-vocabulary correction better at higher cost. Neither fixes noisy advantage estimators.
Results: training Qwen3-30B-A3B-Base on DAPO-Math-17k, BPO averages 50.5% across AIME 2024–2026 — 11 points above GRPO-ClipHigher and 3.1 points above the strongest baseline CISPO.
More from Models
- TeleOCR Trends on Hugging Face: A Qwen2.5-VL-Based Chinese Document OCR Model — XingChen-AGI · 2026-09-28
- Kaggle Game Arena: Evaluating LLMs via Head-to-Head Chess, Poker, and Werewolf — kaggle · 2026-09-28
- Perplexity CEO: still using sol 6 for knowledge work — cheap, fast, great compaction — gabriel1 · 2026-09-28
- NerfBench's First Results Find No Nerf: Claude Opus 5.5 Dips Just 0.8% vs Launch — alejandroll10 · 2026-09-28
- Most humans can read this image instantly — most AI vision models can't — JeremyNguyenPhD · 2026-09-28
- AI-generated 7-minute SQLite repo explainer stuns with coherent code walkthrough — deedydas · 2026-09-28