Score Centering is secretly a STE: new fix targets LLM RL training instability

brandondamos · x · 2026-09-21

@mryabinin traces LLM RL instability — when training and sampling policies differ — to first principles, proposing Score Centering to directly cancel it, simpler and compatible with importance sampling. @YouJiacheng notes the trick is secretly a straight-through estimator: zp + sg(zq - zp).

Related event: RL score centering is really a straight-through estimator, research shows(2 posts)→

Original post →

More from Research

Research channel →