A Handbook on Speculative Decoding: From Serial Generation to min(1,p/q)
techNmak · x · 2026-10-03
The author put together a handbook on how speculative decoding actually works, starting with why autoregressive generation is serial and what changes when a cheaper proposer model drafts several future tokens for the target model to verify together.
It works through greedy verification, exact speculative sampling, the min(1, p/q) acceptance rule, residual correction, acceptance rate, and expected tokens per verification — a systematic resource for understanding this inference acceleration technique.
Related event: Engineer Releases Systematic Handbook on Speculative Decoding(2 posts)→
More from Infra
- Cactus releases Whistle: a 16.9MB speech-to-text model that beats Whisper base on CPU with 6x speed — ycombinator · 2026-10-03
- Cohere and vLLM co-host Toronto meetup on open weights and inference — cohere · 2026-10-03
- TSMC evaluating multi-billion dollar Texas fab campus amid surging US demand from Nvidia, Apple — Beth_Kindig · 2026-10-03
- Discrete diffusion delivers provably lossless LLM inference speedups, drop-in for training — Cohere · 2026-10-03
- Runner pays $4,500 for a third GPU to run near-frontier models locally — marian_nmt · 2026-10-03
- 90% of AI chip value accrues to incumbents, no OpenAI equivalent for years — menhguin · 2026-10-03