RL score centering is really a straight-through estimator, research shows

Researchers analyze RL training instability from first principles, and a derivation shows score centering in policy gradients is essentially a straight-through estimator, explaining why common fixes only sidestep the root cause.

2026-09-20 ~ 2026-09-21 · 2 related posts