Matryoshka Model Suites Cut Compute 36%; Author Mocks Reviewer Culture
Yoav Artzi and Nathan Godey's new paper introduces Matryoshka Language Model Suites, saving about 36% compute via nested training and speeding up inference 26% with speculative decoding. Artzi also mockingly mimicked reviewers who dismiss PPL metrics in the era of trillion-token models.
2026-09-25 ~ 2026-09-25 · 2 related posts
- "Reject, go get more compute": mocking reviewers who dismiss small-scale NLP work — yoavartzi · 2026-09-25
- Matryoshka LM suites: nested training cuts suite compute by 36%, speeds speculative decoding 14-26% — yoavartzi · 2026-09-25