Liquid AI's DSpark speculative decoding makes LFM2.5-VL-3B up to 3.13x faster
helloiamleonie · x · 2026-09-24
Liquid AI released LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model (279.5M params) for its vision-language model LFM2.5-VL-3B. A lightweight drafter proposes tokens ahead and the target verifies them in one pass, leaving output unchanged. Benchmarks: up to 3.13x faster decoding (2.62x end-to-end) with MLX on M5 Max, 2.14x with llama.cpp on M3 Ultra, and 2.66x with SGLang on H100. Open-sourced on Hugging Face with GGUF variants; greedy decoding remains exact.
More from Infra
- Celesto Launches Real Computers for AI Agents With 500ms-Boot MicroVMs — aniketmaurya · 2026-09-24
- Report: Google to Launch Space TPUs on Falcon 9 Next Week to Test Orbital AI Data Centers — ns123abc · 2026-09-24
- Researchers exploited Cloudflare Containers flaw to read other tenants' residual disk data — matthew_d_green · 2026-09-24
- IIT Delhi Develops India's First Indigenous Micro GPU — Paimaamu · 2026-09-24
- Marvell and GlobalFoundries sign multi-year deal to expand silicon germanium capacity — Beth_Kindig · 2026-09-24
- LLM Compressor v0.14.0 makes GPTQ quantization up to 30x faster with new Triton kernel — vllm_project · 2026-09-24