Liquid AI Releases DSpark Draft Models for Up to 3.18x Faster Decoding

JosephJacks_ · x · 2026-08-21

Liquid AI has released DSpark draft models (300M params) for the LFM2.5 series, utilizing speculative decoding. By having a small model propose candidates and a large model verify them, it achieves up to 3.18x decoding speedup on H100 (MATH500) and 2.87x on M4 Max. The technique ensures output remains identical under greedy decoding without compromising accuracy, and is now supported by llama.cpp and SGLang.

Related event: Liquid AI Releases DSpark Draft Model, Accelerating LFM2.5 Decoding by up to 3.18x(5 posts)→

Original post →

More from Infra

Infra channel →