Matryoshka LM suites: nested training cuts suite compute by 36%, speeds speculative decoding 14-26%
yoavartzi · x · 2026-09-25
Yoav Artzi shares his new paper Matryoshka Language Model Suites with Nathan Godey (arXiv:2608.09703), while mocking reviewers who dismiss results over PPL numbers and demand more compute.
Key points:
- Classically, model suites require training and serving each size separately. This framework stacks sub-models of increasing size into a single nested architecture trained end-to-end.
- Enables low-cost distillation from the largest to all smaller sub-models at every training step; the draft model lives inside the verifier, suiting speculative decoding.
- Validated with a 500M/1.5B/3B suite: matches independently trained baselines on benchmarks and perplexities while using 36% less training compute and improving speculative decoding throughput by 14-26%.
- Includes ablations guiding the design of strong Matryoshka LM suites.
Related event: Matryoshka Model Suites Cut Compute 36%; Author Mocks Reviewer Culture(2 posts)→
More from Infra
- AI energy startup Parallax launches with $117m from Founders Fund, Lux, Greylock and others — graceisford · 2026-09-25
- AMD to present MXFP8 pretraining scaling on 1K+ MI355X GPUs at PyTorchCon 2026 — PyTorch · 2026-09-25
- Google's Project Suncatcher puts TPU clusters into orbit for space-based ML infra — rseroter · 2026-09-25
- Nebius/WEKA benchmark: shared KV cache lifts agentic inference throughput 2.4x with 93% hit rate — AccBalanced · 2026-09-25
- Burkov's TP Weekly #179: GPU rent vs buy, llm-d serving 753B model at 5-10x lower cost — burkov · 2026-09-25
- YuE2 music model with full CoT runs on iPhone in just 1.7GB of memory — Acceptable-Cycle4645 · 2026-09-25