Score Centering tackles the root cause of RL instability when train and sample policies differ
_AndrewZhao · x · 2026-09-19
- LLM RL is notoriously unstable when the training and sampling policies use different numerics; standard fixes (matching numerics, importance sampling) only work around the problem.
- The thread derives the root cause from first principles and proposes directly cancelling it.
- Score Centering is simple to implement, competitive on its own, and compatible with existing approaches.
Related event: Together AI's Score Centering Stabilizes Off-Policy RL for LLMs(3 posts)→
More from Research
- Google open-sources Fuse, a multi-agent framework for verifiable social reasoning in LLMs — google · 2026-09-19
- Study finds scientific beauty and impact are surprisingly correlated — jacobkimmel · 2026-09-19
- Trolley problem: Jev-style API turns jina-reranker-v3.5 into a ruthless decision engine — gaganghotra_ · 2026-09-19
- Probe guidance steers continuous diffusion LMs with just 1-3% extra inference compute — itsbautistam · 2026-09-19
- LLM Analysis of Wikipedia Battle Pages Crowns Napoleon the GOAT With 16.7 Career WAR — ctjlewis · 2026-09-19
- MIT paper tracks the birth, life and death of 108k open-weight AI models on Hugging Face — asusarla · 2026-09-19