Liquid AI ships DSpark draft model, up to 3.13x faster decoding for LFM2.5-VL-3B
helloiamleonie · x · 2026-09-25
Liquid AI released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, bringing speculative decoding: a lightweight drafter proposes tokens ahead and the target model verifies them in one pass, with no change in output quality. Benchmarks across six VL task categories (batch 1, temp 0) show up to 3.13x faster decoding on MLX/M5 Max (2.62x end-to-end), 2.14x on llama.cpp/M3 Ultra, and 2.66x on SGLang/H100. Weights are on Hugging Face with day-0 MLX-VLM support.
Related event: Liquid AI's DSpark Speculative Decoding Speeds Up VL Model 3.13x(4 posts)→
More from Infra
- Ubuntu moves to weekly kernel updates as AI finds Linux bugs faster than Canonical can patch — AIFlow_ML · 2026-09-25
- Wasmer runs a real PostgreSQL 18.4 server on iOS and in the browser via WebAssembly — jedisct1 · 2026-09-25
- Inference startup Jatevo returns: 124B tokens, 1.66M requests, $147K of inference delivered — toptickcrypto · 2026-09-25
- Calibration-Free Quantization Method TQ Open-Sourced, Hits 92.4% Top-1 on Qwen 27B 4-bit — textclf · 2026-09-25
- Agentic AI changes the CPU-to-GPU ratio: 5% CPU allocation cuts token cost ~3.7% — BenBajarin · 2026-09-25
- PreFT Paper Accepted at NeurIPS: Prefill-Only LoRA Adapters Speed Up Multi-Adapter Serving — aryaman2020 · 2026-09-25