Researchers Uncover 'Value Flattening' in PPO Critics — Sparse Supervision on 3 States Fixes It

Shanghai-AI-Laboratory · hf · 2026-09-17

Researchers identify a systematic failure mode in PPO critics for RL-based LLM training, dubbed Value Flattening: true state values estimated from multiple Monte Carlo continuations shift sharply across intermediate states while critic predictions remain flat. The effect reproduces in FrozenLake and worsens as state spaces grow.

Experiments on Qwen3-Base show SP³O with just three supervised states per response mitigates Value Flattening and consistently improves the learned policy across model sizes and evaluation suites.

Original post →

More from Research

Research channel →