Liquid AI Releases DSpark Draft Models for Up to 3.18x Faster Decoding
JosephJacks_ · x · 2026-08-21
Liquid AI has released DSpark draft models (300M params) for the LFM2.5 series, utilizing speculative decoding. By having a small model propose candidates and a large model verify them, it achieves up to 3.18x decoding speedup on H100 (MATH500) and 2.87x on M4 Max. The technique ensures output remains identical under greedy decoding without compromising accuracy, and is now supported by llama.cpp and SGLang.
More from Infra
- Cloudera Boosts Spark 4.1 Speed 4x With NVIDIA GPU Acceleration — shashib · 2026-08-21
- Space datacenters won't escape pushback: States will regulate rocket launches — wordgrammer · 2026-08-21
- Solo dev hits $6M run rate with inference provider Morph, aims for $10B valuation with 10 people — ycombinator · 2026-08-21
- Could a Data Center Replace Your Town's Property Taxes? — inductionheads · 2026-08-21
- Americans Prefer Coal Plants Over Data Centers: Study — Polymarket · 2026-08-21
- OpenBMB releases Ultra-FineWeb-L1, a 1T+ token high-quality web dataset — zibuyu9 · 2026-08-21