ROSS relearns from stale rollouts with selective supervision, SWE-bench +4.2 pts
Zhiwei Zhang · hf · 2026-09-30
Self-generated rollouts from LLM post-training are usually discarded once the policy advances, yet they still hold reusable behavioral experience—mixed with mistakes and redundant actions that shouldn't be imitated.
- ROSS keeps full historical trajectories as context but applies loss only to selected model-generated continuations, enabling fine-grained selective supervision
- Consistently improves upstream checkpoints across domain RL, multi-teacher on-policy distillation, and agentic RL, beating baselines on math, code, instruction following, and software engineering
- On Qwen3.6-35B-A3B, six-benchmark MOPD average rises from 58.40% to 62.20%; SWE-bench Verified from 64.20% to 68.40%
- Takeaway: behavioral experience from self-rollout training can be harvested via offline SFT without additional policy rollouts
More from Research
- Researcher predicts AI labs will soon pivot from math conjectures to materials and drug discovery — tak3sh8 · 2026-09-30
- Video lecture series by Stephen Wright, Yousef Saad and Peter Bartlett on ML optimization now available — caglar_ee · 2026-09-30
- Arbor: open-source framework for AI agents doing autonomous long-horizon research — burkov · 2026-09-30
- Physics-aware losses keep grain boundaries real when AI generates alloy microstructures — bravo_abad · 2026-09-30
- SOSP26 Paper YoloFS Targets Agent Filesystem Misuse, Built From 290 Real Incident Reports — tianyin_xu · 2026-09-30
- Agents can delete their own logs: Claude Code, Codex, others fail trace integrity, paper finds — maksym_andr · 2026-09-30