Nebius says SlimSpec speeds speculative decoding 8–9% without shrinking the vocabulary

Arindam_1729 · x · 2026-07-21

Nebius’ Token Factory says its new SlimSpec speculative decoding approach improves serving speed without reducing the vocabulary.

According to the post, SlimSpec cuts LM-head cost by 4–5× while preserving the full vocabulary, and delivers up to 8–9% higher end-to-end speedup versus baselines. The claim is framed as a production-throughput improvement: more useful tokens per second and better serving efficiency.

The image highlights the core idea as “faster speculative decoding without cutting the vocabulary,” positioning it as an infra optimization for real deployment rather than a model-quality change.

Original post →

More from Infra

Infra channel →