Study notes: how speculative decoding accelerates LLM inference without quality loss

helloiamleonie · x · 2026-08-24

Leonie Monigatti published study notes on speculative decoding, the inference optimization technique from DeepMind and Google (2023).

Original post →

More from Infra

Infra channel →