SIS Brings Reused Tokens Closer to On-Policy

A new paper introduces SIS to address the off-policy mismatch caused by reusing rollouts in LLM RL post-training. Instead of only clipping global ratios, it applies token-level importance correction to pull samples closer to on-policy, aiming to reduce re-sampling cost while improving training stability.

2026-07-11 ~ 2026-07-12 · 2 related posts