SFT then RL doesn't fix agent looping: 29% of runs hit turn cap vs 0% for RL alone
VikParuchuri · x · 2026-10-06
Datalab (Vik Paruchuri's team) published a writeup on the agent tool-loop problem, noting looping has been observed before (Liquid AI, and apapiu's work on tool loops in document agents and SFT/RL interaction).
Key findings:
- SFT followed by RL does not fix looping: at temperature 0, 29% of SFT-then-RL runs loop until the turn cap, versus 0% for RL alone.
- The decisive factor appears to be on-policy data — on-policy distillation also loops far less than SFT.
A practical takeaway for agent training teams: whether data is on-policy may matter more for avoiding stuck loops than simply adding an RL stage after SFT.
Related event: Pure RL Post-Training Eliminates Agent Tool-Call Loops, Datalab Finds(3 posts)→
More from coding & agent
- Reflection AI debuts Beam: open agentic model with 501B total, 23B active params — bigblueboo · 2026-10-06
- Better prompting fixed my LLM-built Flappy Bird: full inputs beat images — maximelabonne · 2026-10-06
- Grok Bot announces six live workshops in two weeks, covering marketing, sales and coding — XFreeze · 2026-10-06
- Nous Research publishes an agent manifesto: your model, your memory, your hardware — NousResearch · 2026-10-06
- Guide: Installing the full Google Antigravity Suite (IDE + Hub + CLI) on Linux — CommunicationMean385 · 2026-10-06
- I tested 20+ ways to make a cheap coding model act like an expensive one — KangarooAnxious9394 · 2026-10-06