ActFirst-OPD trains multi-turn agents up to 4.9x faster by acting before reasoning
SouthernUniversityofScienceandTechnology · hf · 2026-09-30
SUSTech proposes ActFirst-OPD, a training framework that decouples environment interaction from full response generation in on-policy distillation (OPD) for multi-turn agents. Key ideas:
- Standard think-then-act rollouts delay environment transitions with lengthy reasoning before each short action. ActFirst flips the order: the student acts first via reference-conditioned inverse dynamics, and falls back to autonomous next-action prediction when transitions deviate from the reference trajectory.
- Full think-then-act responses are then generated asynchronously from collected contexts for token-level teacher supervision.
Results with Qwen3 students (0.6B/1.7B/4B): average wall-clock speedups of 2.3x on ALFWorld, 1.8x on WebShop, and 4.9x on ScienceWorld over vanilla OPD, while matching or beating OPD baselines in 8 of 9 benchmark-model settings on task success rate. Shows reasoning need not block acting in multi-turn agent distillation.
More from coding & agent
- Cursor + data export makes debugging GA4/GSC traffic spikes a 10-minute job — gaganghotra_ · 2026-09-30
- Team builds product-ops agent that learns from past incidents via Hindsight memory — Aggravating_Ice8404 · 2026-09-30
- Intern-Decision open models (0.8B–4B) beat Jev on multimodal decision benchmarks — max_paperclips · 2026-09-30
- xiaohu shows the computer assigned to his AI agent: browser preinstalled and it can make calls — xiaohu · 2026-09-30
- Editable artifacts may beat screenshots for testing agents' visual understanding — OliviaYii · 2026-09-30
- Open-source setup adds multiple providers to Codex desktop via a local Responses proxy — LinkSudah · 2026-09-30