Pure RL Post-Training Eliminates Agent Tool-Call Loops, Datalab Finds

Datalab found that SFT-only post-training caused agents to loop on tool calls 92% of the time, while pure RL training eliminated loops entirely; even SFT-then-RL left 29% of runs hitting step limits, pointing to on-policy data as the key.

2026-10-06 ~ 2026-10-06 · 3 related posts