EasyPPO stabilizes the critic for LLM RL, +14.89% over PPO on coding
Xuanyi Zhou · hf · 2026-09-30
EasyPPO identifies and fixes two critic failure modes that destabilize PPO in LLM reinforcement learning.
- Failure mode 1: filtering truncated rollouts from both actor and critic shifts the objective to reward conditioned on completion, letting truncation increase even as conditional reward improves.
- Failure mode 2: heterogeneous return noise lets high-variance prompts dominate critic updates in finite batches.
- Fixes: actor-only overlong filtering, noise-normalized critic regression (weighting each prompt's critic loss by inverse return std), and moderately smaller critic mini-batches.
- Results: stable across the full training horizon on FrontierCS coding, AIME24 math, and Search-R1 multi-turn search, consistently beating vanilla PPO, VAPO, and HL-Gauss PPO — with relative gains of 14.89%, 2.28%, and 9.47% respectively.
Related event: EasyPPO Stabilizes PPO for LLM Post-Training by Freezing the Critic(2 posts)→
More from Research
- Government weather model WRF ported to GPUs, running an order of magnitude faster with 250m fog forecasts — Scobleizer · 2026-09-30
- François Fleuret: AI math will dwarf human math, focus on lean proofs not explainability — francoisfleuret · 2026-09-30
- SortedRL: Microsoft Research tackles 70-74% GPU idle time in LLM reinforcement learning — burkov · 2026-09-30
- UMass professor Luc Rey-Bellet's stochastic processes lecture notes on Markov chains and MCMC — michaelchchoi · 2026-09-30
- VoxMem benchmark: none of 15 audio LLMs top 40% on spoken multi-session memory — unimelb-hf · 2026-09-30
- EpiCon: shared multimodal memory lifts agent scores 1.7-4.9 points across 11 benchmarks — Ziyun Zeng · 2026-09-30