Together AI paper: score centering cancels drift in off-policy RL under train-inference mismatch
PandaAshwinee · x · 2026-09-18
A new paper from Together AI researchers (Martin Marek & Max Ryabinin) targets RL instability caused by train-inference mismatch (TIM).
- The authors show instability stems primarily from drift: a persistent bias between training and inference engines that accumulates every step
- They derive an additive "score centering" correction that cancels this drift
- Tested on models from 0.6B to 30B parameters: score centering alone matches or beats importance-sampling-based methods, with the gap widening as mismatch grows
- Being additive, it composes with importance sampling, outperforming pure IS baselines in staleness experiments
Related event: Together AI Proposes Score Centering to Fix RL Train-Inference Mismatch(2 posts)→
More from Research
- Turning Yang-Mills existence and mass gap into a formal conjecture is AI math's ultimate test — geoffreyirving · 2026-09-18
- CoRL 2026 SPIN workshop pits robots against a child in live dexterity challenge — berkeley_ai · 2026-09-18
- That viral "air-gapped data exfiltration" paper only read temperature over a 4cm gap — basedjensen · 2026-09-18
- An AI forecaster has won the seasonal Metaculus Cup for the first time — NathanpmYoung · 2026-09-18
- New video: weather forecasting with neural networks explained — ariG23498 · 2026-09-18
- First LLM runs entirely on-chain: weights, activations, attention all in SVM transactions — AccBalanced · 2026-09-18