Liquid AI ships DSpark drafter, up to 3.13x faster decoding for LFM2.5-VL-3B
JosephJacks_ · x · 2026-09-25
Liquid AI released an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to its vision-language models: a lightweight drafter proposes multiple tokens and the target model verifies them in one pass, without quality loss. Benchmarks (batch size 1, temperature 0): MLX on M5 Max up to 3.13x faster decoding (2.62x end-to-end); llama.cpp on M3 Ultra 2.14x (1.77x e2e); SGLang on H100 2.66x (2.27x e2e), collected via its Pipette benchmarking infra.
Related event: Liquid AI's DSpark Speculative Decoding Speeds Up VL Model 3.13x(4 posts)→
More from Infra
- Moody's: Big five hyperscalers' future commitments hit $2.8T, up from $350B in 2023 — rohanpaul_ai · 2026-09-25
- Ubuntu moves to weekly kernel updates as AI finds Linux bugs faster than Canonical can patch — AIFlow_ML · 2026-09-25
- Wasmer runs a real PostgreSQL 18.4 server on iOS and in the browser via WebAssembly — jedisct1 · 2026-09-25
- Inference startup Jatevo returns: 124B tokens, 1.66M requests, $147K of inference delivered — toptickcrypto · 2026-09-25
- Calibration-Free Quantization Method TQ Open-Sourced, Hits 92.4% Top-1 on Qwen 27B 4-bit — textclf · 2026-09-25
- Agentic AI changes the CPU-to-GPU ratio: 5% CPU allocation cuts token cost ~3.7% — BenBajarin · 2026-09-25