SIS Brings Off-Policy Tokens Back On-Policy
burny_tech · x · 2026-07-12
- This paper introduces SIS to solve the "reusing rollouts after the policy has changed" problem in LLM reinforcement learning.
- Instead of merely clipping the off-policy ratio, it performs per-token detection: if the current model would still sample the token, it's treated as on-policy; otherwise, standard correction applies.
- The author claims this reintegrates tokens previously deemed off-policy back into on-policy training.
- Integrating SIS into GRPO, DAPO, and GSPO yielded improvements in math tasks and agentic search, while stabilizing training under stale rollout and MoE mismatch scenarios.
Related event: SIS Brings Reused Tokens Closer to On-Policy(2 posts)→
More from Research
- Cold Spring Harbor Asia sets a genome biology conference in Suzhou for Oct. 12–16 — jmuiuc · 2026-07-21
- A clean counterexample shows a map can be locally diffeomorphic yet globally fold — Algomancer · 2026-07-21
- Xiaohongshu’s dots-note-3.0 gets a perfect IMO score and becomes the world’s second gold model — 量子位 · 2026-07-21
- Statistical theory paper studies how fast signatures learn in path regression — chaumian · 2026-07-21
- PROWL uses a world model to keep Minecraft agents exploring after failures — nathanbenaich · 2026-07-21
- LeRobot v0.6.0 adds end-to-end 3D depth training data for robots — RemiCadene · 2026-07-21