Two-year-old GPO paper sits at the frontier of RL scaling, recipes now public

QuanquanGu · x · 2026-09-19

Yifan Zhang highlights that GPO, a paper published two years ago, is still at the frontier of RL scaling, declaring that "frontier RL recipes have been revealed" and pointing to three related works: GPO, RPG, and BPO. The thread traces how early RL methods anticipate today's reasoning-model training pipelines.

Original post →

More from Research

Research channel →