Liquid AI Releases DSpark: Speculative Decoding Up to 3.18x Faster
helloiamleonie · x · 2026-08-21
Liquid AI released DSpark draft models for the LFM2.5 series. This technique uses a lightweight draft model to propose candidate tokens, which are then verified by the target model in a single forward pass. It achieves significant decoding speedups with minimal memory overhead and no change in output quality. Benchmarks show up to 3.18x throughput on H100 (MATH500) and 2.87x on MacBook Pro M4 Max (HumanEval).
More from Infra
- Cursor's Git Storage System: S3 as Source of Truth, Local Disk as Cache — xennygrimmato_ · 2026-08-21
- Google Antigravity Expands to VS Code, Zed, JetBrains with Enterprise Controls — rseroter · 2026-08-21
- Cursor Deep Dive: Engineering Challenges of Hosting Git at Scale — JeremyCMorgan · 2026-08-21
- Unsloth Desktop Update: Auto Compaction and LAN Remote Access — danielhanchen · 2026-08-21
- AMD ROCm 10.1 fixes major issues: LLaMA.cpp runs flawlessly on RDNA2 — smellof · 2026-08-21
- RAG vs CAG: KV Cache Cuts LLM Costs by 90% — blaizedsouza · 2026-08-21