EasyPPO: just fix the critic — stable PPO for LLM post-training with zero training collapse
teortaxesTex · x · 2026-09-30
- The idea: EasyPPO brings PPO back to LLM post-training by changing exactly one thing — fix the critic. No new actor loss, no new policy algorithm, actor updates unchanged. Across all experiments it shows zero training collapse while delivering better results.
- Why PPO again: in a quote-thread, wenhaocha1 argues GRPO is hard to scale because it burns rollouts on every task. For math/coding an extra rollout is just a few more compiles and checks, but for RSI a rollout is a training run costing thousands of dollars, and in biology you can't run parallel identical embryos. The goal: get far more out of each rollout.
- The team shared early work on keeping PPO training stable and open-sourced it.
Related event: EasyPPO Stabilizes PPO for LLM Post-Training by Freezing the Critic(2 posts)→
More from Infra
- AAOI burnt capex on US vertical integration, CW laser yield still poor — jwt0625 · 2026-09-30
- PrismQuant: null-space rotations make INT4 near-lossless, only 0.22pp below FP16 on Llama-70B — NanyangTechnologicalUniversity · 2026-09-30
- AI data centers may push IG rates past 10% as compute absorbs global capital — sudoraohacker · 2026-09-30
- ByteDance's HELIX Unifies Sequence Retrieval and Feature Interaction, Deployed in TikTok — _reachsumit · 2026-09-30
- DeepSeek's releases reportedly come surprisingly close to a full Ascend stack for frontier training — teortaxesTex · 2026-09-30
- Strata engine runs Qwen3.8 Flash Next on a 12GB laptop at 51 t/s, 1500 t/s prefill — MLDataScientist · 2026-09-30