ROFT: fine-tuning only on self-explanations matches GRPO on SWE-bench without RL
iScienceLuvr · x · 2026-09-29
A new arXiv paper introduces Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure where an agent attempts a task, generates a retrospective explanation of its experience, and is fine-tuned with next-token prediction loss on the explanation tokens alone — no external teacher, no reward-based policy update.
Key results:
- Trained on Qwen3.5-4B using a mix of successful and failed base-model attempts
- Reaches 49.2% on SWE-bench Verified and 26.8% on SWE-bench Pro after 20 updates, no verifier needed
- Beats GRPO's 48.0% / 25.3% after 40 updates, with faster early progress in training time and sampled attempts
- Learns to solve tasks where all 64 sampled base-model attempts failed, showing learning can begin without any initially successful trajectories
More from coding & agent
- Undoing one AI edit without losing the rest: change-record replay with staleness checks — memokris · 2026-09-29
- Hindsight: turning old post-mortems into new system constraints — syedanisafatima · 2026-09-29
- Same Prompt, Two Eras: GPT-5.5 vs Claude Sonnet 4 Flappy Bird Test Shows 1.5 Years of LLM Gains — Dhakkad_Mutthal · 2026-09-29
- Built a Full Halloween Game in Minutes: Code, Art and Music All AI-Generated — iamfakhrealam · 2026-09-29
- Raven 0.2.0 adds self-improving harnesses and shared task graphs for multi-agent workflows — nikola_mr64990 · 2026-09-29
- Anthropic engineer runs 100+ Claude Code agents with Chief and PM agents managing the workflow — anthara_ai · 2026-09-29