Score Centering is secretly a STE: new fix targets LLM RL training instability
brandondamos · x · 2026-09-21
@mryabinin traces LLM RL instability — when training and sampling policies differ — to first principles, proposing Score Centering to directly cancel it, simpler and compatible with importance sampling. @YouJiacheng notes the trick is secretly a straight-through estimator: zp + sg(zq - zp).
Related event: RL score centering is really a straight-through estimator, research shows(2 posts)→
More from Research
- Jev's eval abstraction maps 1:1 to autorubric paper from 8 months ago, researcher finds — deliprao · 2026-09-21
- Researcher Says TypeSafe's Jev Mirrors His Autorubric LLM Eval Framework From 8 Months Ago — deliprao · 2026-09-21
- As ICLR 2027 Tops 60K Submissions, Researcher Proposes 3-4 Paper Cap per Author — ziv_ravid · 2026-09-21
- ICLR 2027 Hits 60K+ Submissions; Researcher Proposes Paper Caps, Forced Reproducibility — ziv_ravid · 2026-09-21
- No, Laya isn't capped at 512 tokens — it's ModernBERT with 8192-token configs — antoine_chaffin · 2026-09-21
- Kev open-source decision models scale to 0.6B/4B/8B, trainable in 40 min on one H100 — TheMoonMidas · 2026-09-21