Deep-Dive Speculative Decoding Blog Incoming: Drafter Training to vLLM Serving
auto_grad_ · x · 2026-09-10
The author, whose speculative decoding experiments have generally panned out, is releasing a highly technical blog covering: why it's needed from an arithmetic-intensity perspective, picking a drafter setup probabilistically, structuring verification for max efficiency, training drafters to align with the target distribution (SFT to OPD), and serving efficiently with vLLM.
More from Infra
- turbovec: Rust vector index fits 10M document vectors in 4GB RAM and outpaces FAISS — tom_doerr · 2026-09-10
- LM Studio 0.4.24 adds advanced llama.cpp argument overrides for GGUF model loading — solyarisoftware · 2026-09-10
- tszzl wraps up: efficiency gains only amplify hunger for hardware — tszzl · 2026-09-10
- Trimming MTP draft vocab to 47k boosts DGX Spark code decoding by 21.5% on same hardware — MaziyarPanahi · 2026-09-10
- Is 5 tokens/s usable for local LLMs? Redditor runs 27B model off an iGPU — Zombiecidialfreak · 2026-09-10
- KV cache exposes agent economics: devs pay big for context re-reads that cost providers nothing — hackgoofer · 2026-09-10