mlx-vlm Integrates LFM2.5 DSpark for 3.7x Speedup on M5 Max
helloiamleonie · x · 2026-08-22
mlx-vlm added support for Liquid AI's LFM2.5 DSpark in v0.6.16. Using exact speculative decoding, it achieves up to 3.7x faster generation on M5 Max with zero output drift. The implementation includes shared Metal GEMV kernels and a hybrid-capture rollback mechanism to ensure exact parity for both dense and MoE models.
More from Infra
- MCP Isn't Replacing APIs: It's Changing Who APIs Are Designed For — kush_patil · 2026-08-22
- Data Center Opposition Surged from 42 to 75 Percent in One Year — The Decoder · 2026-08-22
- Qwen3.8-27B gets DFlash2 speculative-decoding GGUF release for llama.cpp — incoai · 2026-08-22
- Woolly post-trains Qwen3-8B for 2–3× faster math & code decoding — bosmeny · 2026-08-22
- Opinion: States Banning Data Centers Face 20 Years of Economic Depression — GabGarrett · 2026-08-22
- Vercel Fixes TLS Fragmentation Issue, All Websites Back Online — uwukko · 2026-08-22