New handbook explains speculative decoding: why acceptance rate α=0.8 still yields only 1.87× speedup
techNmak · x · 2026-10-03
Engineer techNmak released a comprehensive handbook on speculative decoding, starting from why autoregressive generation is serial and how a cheaper proposer drafting multiple tokens for joint verification by the target model changes the game.
The handbook covers greedy verification, exact speculative sampling with the min(1, p/q) acceptance rule, residual correction, acceptance rate, expected tokens per verification, draft-length vs. wall-time trade-offs, KV-cache commit and rollback, tokenizer compatibility, self-speculation, Medusa, SpecInfer, EAGLE, MTP, prompt lookup, dynamic speculation, and current serving implementations — with worked examples and citations throughout.
A key point: acceptance rate is not speedup. In one example, α = 0.8 and γ = 4 yields 3.3616 expected tokens per target verification, but when draft cost rises from 0.05 to 0.20 of a target step, the predicted speedup drops from 2.80× to 1.87×.
Related event: Engineer Releases Systematic Handbook on Speculative Decoding(2 posts)→
More from Infra
- Suhail hails the start of the Vera Rubin era as NVIDIA's next-gen GPU boots up — rickasaurus · 2026-10-03
- GPU-backed loans put GPU earning power and collateral value under scrutiny — AnneliesGamble · 2026-10-03
- Google accused of illegally bulldozing 300 million sq. meters of Finnish forest for AI data centers — Polymarket · 2026-10-03
- SkyRL v0.4 trains 1T-param Kimi K2.7 with RL on just 16 B300 GPUs — casper_hansen_ · 2026-10-03
- BIS report: 55% of AI investment is circular, echoing Lucent-Nortel era risks — rohanpaul_ai · 2026-10-03
- RAM prices may 10-20x next year and double again by 2028, predicts analyst — Kyrannio · 2026-10-03