mlx-vlm Multi-Turn Latency Drops to 0.25s

WesEklund · x · 2026-07-11

The author details their contribution to an open-source local AI project, reducing the multi-turn agent warm latency for mlx-vlm from roughly 72 seconds to about 0.25 seconds, marking a 290x speedup.

They explain two existing optimizations:

Previously, these two conflicted—enabling KV quantization essentially broke prefix caching. This limitation has now been resolved, allowing developers to leverage both multi-turn speed and memory efficiency simultaneously. The author mentions the new version will be released "very soon."

Related event: Cache Fix Reduces mlx-vlm Agent Latency by 290x(4 posts)→

Original post →

More from coding & agent

coding & agent channel →