SMRC-SD: Solving State-Reference Mismatch in Multi-Turn Agent Distillation

Junzhuo Liu · hf · 2026-08-10

When using successful trajectories for policy distillation in multi-turn agents, the student's execution state often mismatches the reference trajectory. To address this, researchers introduced SMRC-SD (State-Matched Routing and Contextualized Self-Distillation).

Core mechanisms:

Across ALFWorld and WebShop, this method significantly boosts task success rates. Using the Qwen3-1.7B model, success rates improved from 0.746 to 0.865 on ALFWorld and from 0.574 to 0.693 on WebShop.

Original post →

More from coding & agent

coding & agent channel →