LFM2.5-DSpark offers up to 3.2x faster inference
pmttyji · reddit · 2026-08-21
LFM2.5-DSpark models are now available with up to 3.2x faster inference using speculative decoding. Official GGUF files for 1.2B, 2.6B, and 8B variants have been released on Hugging Face. The author advises against using testing versions from the PR and asks for advice on running these on mobile devices.
More from Infra
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24
- LifeOS: A Local, Voice-Driven Personal Organizer — Extension-Bid-639 · 2026-08-24