mlx-vlm Integrates LFM2.5 DSpark for 3.7x Speedup on M5 Max

helloiamleonie · x · 2026-08-22

mlx-vlm added support for Liquid AI's LFM2.5 DSpark in v0.6.16. Using exact speculative decoding, it achieves up to 3.7x faster generation on M5 Max with zero output drift. The implementation includes shared Metal GEMV kernels and a hybrid-capture rollback mechanism to ensure exact parity for both dense and MoE models.

Original post →

More from Infra

Infra channel →