ROSS relearns from stale rollouts with selective supervision, SWE-bench +4.2 pts

Zhiwei Zhang · hf · 2026-09-30

Self-generated rollouts from LLM post-training are usually discarded once the policy advances, yet they still hold reusable behavioral experience—mixed with mistakes and redundant actions that shouldn't be imitated.

Original post →

More from Research

Research channel →