SIS Brings Reused Tokens Closer to On-Policy
A new paper introduces SIS to address the off-policy mismatch caused by reusing rollouts in LLM RL post-training. Instead of only clipping global ratios, it applies token-level importance correction to pull samples closer to on-policy, aiming to reduce re-sampling cost while improving training stability.
2026-07-11 ~ 2026-07-12 · 2 related posts
- SIS: Turning Off-Policy Tokens Back to On-Policy — 青稞AI · 2026-07-11
- SIS Brings Off-Policy Tokens Back On-Policy — burny_tech · 2026-07-12