Liquid AI Releases DSpark: Speculative Decoding Up to 3.18x Faster

helloiamleonie · x · 2026-08-21

Liquid AI released DSpark draft models for the LFM2.5 series. This technique uses a lightweight draft model to propose candidate tokens, which are then verified by the target model in a single forward pass. It achieves significant decoding speedups with minimal memory overhead and no change in output quality. Benchmarks show up to 3.18x throughput on H100 (MATH500) and 2.87x on MacBook Pro M4 Max (HumanEval).

Original post →

More from Infra

Infra channel →