slime merges Score Centering, a first-principles fix for RL instability when train and sampling policies diverge

hsu_byron · x · 2026-09-23

The open-source RL training framework slime (THUDM, 8.5k stars) has merged Score Centering support, enabled via --use-score-centering.

Original post →

More from coding & agent

coding & agent channel →