ActFirst-OPD trains multi-turn agents up to 4.9x faster by acting before reasoning

SouthernUniversityofScienceandTechnology · hf · 2026-09-30

SUSTech proposes ActFirst-OPD, a training framework that decouples environment interaction from full response generation in on-policy distillation (OPD) for multi-turn agents. Key ideas:

Results with Qwen3 students (0.6B/1.7B/4B): average wall-clock speedups of 2.3x on ALFWorld, 1.8x on WebShop, and 4.9x on ScienceWorld over vanilla OPD, while matching or beating OPD baselines in 8 of 9 benchmark-model settings on task success rate. Shows reasoning need not block acting in multi-turn agent distillation.

Original post →

More from coding & agent

coding & agent channel →