Pure RL Post-Training Eliminates Agent Tool-Call Loops, Datalab Finds
Datalab found that SFT-only post-training caused agents to loop on tool calls 92% of the time, while pure RL training eliminated loops entirely; even SFT-then-RL left 29% of runs hitting step limits, pointing to on-policy data as the key.
2026-10-06 ~ 2026-10-06 · 3 related posts
- RL Post-Training Eliminates Agent Tool-Call Loops: 92% Loop Rate Drops to 0 — VikParuchuri · 2026-10-06
- RL Alone Eliminates Agent Loops: 0% at Temp 0 vs 92% for SFT, Datalab Finds — VikParuchuri · 2026-10-06
- SFT then RL doesn't fix agent looping: 29% of runs hit turn cap vs 0% for RL alone — VikParuchuri · 2026-10-06