Liquid AI releases DSpark draft models, speeding up decoding by up to 4x
JosephJacks_ · x · 2026-08-21
Liquid AI has released DSpark draft models for the LFM2.5 series, utilizing speculative decoding to significantly accelerate inference speeds.
- Mechanism: A lightweight draft model proposes candidate tokens, and the target model verifies them in a single forward pass, trading minimal memory increase for substantial speedups.
- Performance Data:
- On H100: LFM2.5-8B-A1B on MATH500 increased from 428 to 1362 tok/s (3.18x).
- On-device (M4 Max): LFM2.5-1.2B-Instruct on HumanEval increased from 136 to 389 tok/s (2.87x).
- LFM2.5-2.6B achieved 2.67x speedup on H100.
- Use Cases: Ideal for on-device function calling applications or cloud workloads with critical latency requirements.
More from Models
- ARC Prize Adds Model Comparison, Gemini 3.7 Flash Scores High — mhmazur · 2026-08-21
- Anthropic's Fable Breaks RareBench Record After Relaxing Filters — danielmckinn0n · 2026-08-21
- NVIDIA Explains Omni-Models: Unified Architecture for Text, Images, Audio, Video, and Actions — NVIDIA Developer · 2026-08-21
- Monitors Detect Significant Behavior Shift in Claude Opus — altryne · 2026-08-21
- Users report GPT-4.1 Sol model suddenly became dumb with irrelevant answers — M-M103 · 2026-08-21
- OpenAI pauses largest training run ever as DeepSeek 'cooks again' — Fireship · 2026-08-21