Draft-OPD: On-Policy Distillation Boosts Speculative Decoding
auto_grad_ · x · 2026-08-25
Accepted to #EMNLP2026, this paper addresses the train-serve mismatch in speculative decoding. Traditional draft models are trained via offline SFT on target-generated trajectories, but during inference, they face states induced by their own policy, leading to performance plateaus. The proposed Draft-OPD method uses target-assisted rollouts for stable continuations and replays drafting from verification-exposed error positions. This allows the drafter to learn from target feedback on both accepted and rejected proposals, significantly improving acceptance length and inference speed.
More from Infra
- NVIDIA Vera Rubin benchmarks show 35x cheaper agentic coding tokens — brianryhuang · 2026-08-25
- Dell Partners with NVIDIA and Groq to Deploy 3,400 Tokens/sec Inference Solution — IanAndrewsDC · 2026-08-25
- OpenAI Files Five Pepper-Named Chip Trademarks in a Single Day — AJChadha · 2026-08-25
- Analyst expects TPU shipments to surpass NVIDIA's by 2028 — AccBalanced · 2026-08-25
- Guide: Running Hermes Agent on a Raspberry Pi — LeviTurk · 2026-08-25
- West Virginia targets data centers; proximity to nuclear reactors cited as a key advantage — mimi10v3 · 2026-08-25