Matryoshka Language Model Suites Cuts Training Compute by 36%

willdepue · x · 2026-08-21

This paper proposes a 'Matryoshka' training framework that nests 500M, 1.5B, and 3B models within a single architecture for end-to-end training. This allows smaller models to receive near-free distillation from the largest model and share weights + KV cache for speculative decoding. The suite matches baseline performance while using 36% less training compute and improving speculative decoding throughput by 14-26%.

Related event: Matryoshka LM Suites: Nested Training Cuts Compute by 36%(8 posts)→

Original post →

More from Research

Research channel →