MLX-VLM Adds Disk Offload Support for Qwen4 PLE Tables

EyalToledano · x · 2026-08-29

A PR was merged into MLX-VLM adding pleondisk mode for Qwen4exp models. This feature keeps the n-gram PLE table on disk and reads/dequantizes rows on demand instead of holding them in memory. Similar to MoE offloading, it aims to reduce memory usage for models with large, mostly cold weight structures, optimizing inference efficiency.

Original post →

More from Infra

Infra channel →