SFT then RL doesn't fix agent looping: 29% of runs hit turn cap vs 0% for RL alone

VikParuchuri · x · 2026-10-06

Datalab (Vik Paruchuri's team) published a writeup on the agent tool-loop problem, noting looping has been observed before (Liquid AI, and apapiu's work on tool loops in document agents and SFT/RL interaction).

Key findings:

A practical takeaway for agent training teams: whether data is on-policy may matter more for avoiding stuck loops than simply adding an RL stage after SFT.

Related event: Pure RL Post-Training Eliminates Agent Tool-Call Loops, Datalab Finds(3 posts)→

Original post →

More from coding & agent

coding & agent channel →