Liquid AI Ships DSpark Draft Model, Speeding Up LFM2.5-VL-3B Decoding Up to 3.13x
JosephJacks_ · x · 2026-09-25
Liquid AI released an experimental DSpark draft model bringing speculative decoding to its LFM2.5-VL-3B vision-language model: a lightweight drafter proposes multiple tokens ahead, verified by the target model in a single pass, with no change to output quality.
Benchmarked across six vision-language task categories at batch size 1 and temperature 0:
- MLX on M5 Max: up to 3.13x faster decoding, 2.62x end-to-end
- llama.cpp on M3 Ultra: up to 2.14x faster decoding, 1.77x end-to-end
- SGLang on H100: up to 2.66x faster decoding, 2.27x end-to-end
All results were collected with Pipette, the benchmarking infrastructure behind Liquid AI's public evaluations.
Related event: Liquid AI's DSpark Speculative Decoding Speeds Up VL Model 3.13x(4 posts)→
More from Infra
- Google's Project Suncatcher flies TPU prototype satellite on SpaceX Transporter-18 — Miles_Brundage · 2026-09-25
- YC-backed Isoquant launches GLM-5.3-Flash inference at $0.07/M with 452ms TTFT — ycombinator · 2026-09-25
- Nemotron 3 Speaker Diarization Ported to Apple Silicon via Core ML and MLX — ivan_digital · 2026-09-25
- Strangely, GPU matmuls run faster on 'predictable' data: Horace He explains — goyal__pramod · 2026-09-25
- PyTorch announces ExecuTorch Hackathon in San Francisco, Oct 17-18, with three device tracks — PyTorch · 2026-09-25
- Speculation: GPT-6 Luna/Sol efficiency lean hints at Cerebras 1000 tok/s inference economics — brandon_galang · 2026-09-25