RL Alone Eliminates Agent Loops: 0% at Temp 0 vs 92% for SFT, Datalab Finds

VikParuchuri · x · 2026-10-06

While fine-tuning a small model to drive its Document Agent (scientific PDF → JATS XML, benchmarked on 100 held-out articles), Datalab found the dreaded tool-call looping depends heavily on training method: SFT loops 92% of the time at temp 0, SFT-then-RL still loops 47% (29% hitting the turn cap), while RL from the base model loops 0%. The key appears to be on-policy data — on-policy distillation also loops far less than SFT. Both approaches lift benchmark scores, but only RL consistently fixes looping, a directly actionable finding for anyone posttraining agent models.

Related event: Pure RL Post-Training Eliminates Agent Tool-Call Loops, Datalab Finds(3 posts)→

Original post →

More from coding & agent

coding & agent channel →