Liquid AI ships DSpark drafter, up to 3.13x faster decoding for LFM2.5-VL-3B

JosephJacks_ · x · 2026-09-25

Liquid AI released an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to its vision-language models: a lightweight drafter proposes multiple tokens and the target model verifies them in one pass, without quality loss. Benchmarks (batch size 1, temperature 0): MLX on M5 Max up to 3.13x faster decoding (2.62x end-to-end); llama.cpp on M3 Ultra 2.14x (1.77x e2e); SGLang on H100 2.66x (2.27x e2e), collected via its Pipette benchmarking infra.

Related event: Liquid AI's DSpark Speculative Decoding Speeds Up VL Model 3.13x(4 posts)→

Original post →

More from Infra

Infra channel →