LFM2.5-DSpark offers up to 3.2x faster inference

pmttyji · reddit · 2026-08-21

LFM2.5-DSpark models are now available with up to 3.2x faster inference using speculative decoding. Official GGUF files for 1.2B, 2.6B, and 8B variants have been released on Hugging Face. The author advises against using testing versions from the PR and asks for advice on running these on mobile devices.

Related event: LiquidAI's LFM2.5-DSpark Gets 3.2x Speedup with Speculative Decoding, GGUF Released(2 posts)→

Original post →

More from Infra

Infra channel →