Liquid AI's DSpark speculative decoding makes LFM2.5-VL-3B up to 3.13x faster

helloiamleonie · x · 2026-09-24

Liquid AI released LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model (279.5M params) for its vision-language model LFM2.5-VL-3B. A lightweight drafter proposes tokens ahead and the target verifies them in one pass, leaving output unchanged. Benchmarks: up to 3.13x faster decoding (2.62x end-to-end) with MLX on M5 Max, 2.14x with llama.cpp on M3 Ultra, and 2.66x with SGLang on H100. Open-sourced on Hugging Face with GGUF variants; greedy decoding remains exact.

Original post →

More from Infra

Infra channel →