ROFT: fine-tuning only on self-explanations matches GRPO on SWE-bench without RL

iScienceLuvr · x · 2026-09-29

A new arXiv paper introduces Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure where an agent attempts a task, generates a retrospective explanation of its experience, and is fine-tuned with next-token prediction loss on the explanation tokens alone — no external teacher, no reward-based policy update.

Key results:

Original post →

More from coding & agent

coding & agent channel →