Draft-OPD: On-Policy Distillation Boosts Speculative Decoding

auto_grad_ · x · 2026-08-25

Accepted to #EMNLP2026, this paper addresses the train-serve mismatch in speculative decoding. Traditional draft models are trained via offline SFT on target-generated trajectories, but during inference, they face states induced by their own policy, leading to performance plateaus. The proposed Draft-OPD method uses target-assisted rollouts for stable continuations and replays drafting from verification-exposed error positions. This allows the drafter to learn from target feedback on both accepted and rejected proposals, significantly improving acceptance length and inference speed.

Original post →

More from Infra

Infra channel →