Liquid AI ships DSpark draft model, up to 3.13x faster decoding for LFM2.5-VL-3B

helloiamleonie · x · 2026-09-25

Liquid AI released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, bringing speculative decoding: a lightweight drafter proposes tokens ahead and the target model verifies them in one pass, with no change in output quality. Benchmarks across six VL task categories (batch 1, temp 0) show up to 3.13x faster decoding on MLX/M5 Max (2.62x end-to-end), 2.14x on llama.cpp/M3 Ultra, and 2.66x on SGLang/H100. Weights are on Hugging Face with day-0 MLX-VLM support.

Related event: Liquid AI's DSpark Speculative Decoding Speeds Up VL Model 3.13x(4 posts)→

Original post →

More from Infra

Infra channel →