RL Alone Eliminates Agent Loops: 0% at Temp 0 vs 92% for SFT, Datalab Finds
VikParuchuri · x · 2026-10-06
While fine-tuning a small model to drive its Document Agent (scientific PDF → JATS XML, benchmarked on 100 held-out articles), Datalab found the dreaded tool-call looping depends heavily on training method: SFT loops 92% of the time at temp 0, SFT-then-RL still loops 47% (29% hitting the turn cap), while RL from the base model loops 0%. The key appears to be on-policy data — on-policy distillation also loops far less than SFT. Both approaches lift benchmark scores, but only RL consistently fixes looping, a directly actionable finding for anyone posttraining agent models.
Related event: Pure RL Post-Training Eliminates Agent Tool-Call Loops, Datalab Finds(3 posts)→
More from coding & agent
- Cursor SDK subagents now report back to parent, system prompt becomes replaceable — tetsuoai · 2026-10-06
- Grok Bot announces six live workshops in two weeks, covering marketing, sales and coding — XFreeze · 2026-10-06
- Nous Research publishes an agent manifesto: your model, your memory, your hardware — NousResearch · 2026-10-06
- Guide: Installing the full Google Antigravity Suite (IDE + Hub + CLI) on Linux — CommunicationMean385 · 2026-10-06
- I tested 20+ ways to make a cheap coding model act like an expensive one — KangarooAnxious9394 · 2026-10-06
- Dev Takes All Repos From Agent Readiness Level 1 to 5 in One Week — matanSF · 2026-10-06