RL score centering is really a straight-through estimator, research shows
Researchers analyze RL training instability from first principles, and a derivation shows score centering in policy gradients is essentially a straight-through estimator, explaining why common fixes only sidestep the root cause.
2026-09-20 ~ 2026-09-21 · 2 related posts
- Elegant Derivation Shows Policy Gradient Score Centering Is a Straight-Through Estimator — YouJiacheng · 2026-09-20
- Score Centering is secretly a STE: new fix targets LLM RL training instability — brandondamos · 2026-09-21