New handbook explains speculative decoding: why acceptance rate α=0.8 still yields only 1.87× speedup

techNmak · x · 2026-10-03

Engineer techNmak released a comprehensive handbook on speculative decoding, starting from why autoregressive generation is serial and how a cheaper proposer drafting multiple tokens for joint verification by the target model changes the game.

The handbook covers greedy verification, exact speculative sampling with the min(1, p/q) acceptance rule, residual correction, acceptance rate, expected tokens per verification, draft-length vs. wall-time trade-offs, KV-cache commit and rollback, tokenizer compatibility, self-speculation, Medusa, SpecInfer, EAGLE, MTP, prompt lookup, dynamic speculation, and current serving implementations — with worked examples and citations throughout.

A key point: acceptance rate is not speedup. In one example, α = 0.8 and γ = 4 yields 3.3616 expected tokens per target verification, but when draft cost rises from 0.05 to 0.20 of a target step, the predicted speedup drops from 2.80× to 1.87×.

Related event: Engineer Releases Systematic Handbook on Speculative Decoding(2 posts)→

Original post →

More from Infra

Infra channel →