Matryoshka: Speculative Decoding Speeds Up Inference by 10-30%

nthngdy · x · 2026-08-19

At inference time, speculative decoding is performed on the largest model using smaller submodels as drafts. Nesting allows shared memory and activations, amortizing draft overhead. This achieves 10-30% faster decoding. During training, all submodels are trained together, enabling free student-teacher online distillation during pretraining for better alignment.

Related event: Matryoshka LM Suites: Nested Training Cuts Compute by 36%(8 posts)→

Original post →

More from Infra

Infra channel →