A Handbook on Speculative Decoding: From Serial Generation to min(1,p/q)

techNmak · x · 2026-10-03

The author put together a handbook on how speculative decoding actually works, starting with why autoregressive generation is serial and what changes when a cheaper proposer model drafts several future tokens for the target model to verify together.

It works through greedy verification, exact speculative sampling, the min(1, p/q) acceptance rule, residual correction, acceptance rate, and expected tokens per verification — a systematic resource for understanding this inference acceleration technique.

Related event: Engineer Releases Systematic Handbook on Speculative Decoding(2 posts)→

Original post →

More from Infra

Infra channel →