BPCO paper distills a stable PPO recipe for LLM RL, beating GRPO across scales

max_paperclips · x · 2026-09-03

Researchers introduce Best-Practice Critic Optimization (BPCO), a practical recipe for training critics reliably in LLM RL.

Paper and code are public; the retweeter calls it a must-read for practitioners for its approach to doing science alone.

Original post →

More from Research

Research channel →