Matryoshka LM suites: nested training cuts suite compute by 36%, speeds speculative decoding 14-26%

yoavartzi · x · 2026-09-25

Yoav Artzi shares his new paper Matryoshka Language Model Suites with Nathan Godey (arXiv:2608.09703), while mocking reviewers who dismiss results over PPL numbers and demand more compute.

Key points:

Related event: Matryoshka Model Suites Cut Compute 36%; Author Mocks Reviewer Culture(2 posts)→

Original post →

More from Infra

Infra channel →