EasyPPO Stabilizes PPO for LLM Post-Training by Freezing the Critic
Xuanyi Zhou released EasyPPO, which stabilizes PPO for LLM post-training by simply freezing the critic, addressing two key critic-related failure modes and achieving a 14.89% improvement over PPO on coding tasks.
2026-09-30 ~ 2026-09-30 · 2 related posts
- EasyPPO stabilizes the critic for LLM RL, +14.89% over PPO on coding — Xuanyi Zhou · 2026-09-30
- EasyPPO: just fix the critic — stable PPO for LLM post-training with zero training collapse — teortaxesTex · 2026-09-30