MLX-VLM Adds Disk Offload Support for Qwen4 PLE Tables
EyalToledano · x · 2026-08-29
A PR was merged into MLX-VLM adding pleondisk mode for Qwen4exp models. This feature keeps the n-gram PLE table on disk and reads/dequantizes rows on demand instead of holding them in memory. Similar to MoE offloading, it aims to reduce memory usage for models with large, mostly cold weight structures, optimizing inference efficiency.
More from Infra
- Ollama's New Claude Toggle Lets You Run Local Models Inside Claude Desktop — dr_cintas · 2026-08-29
- AI datacenters face visceral physical backlash as expansion meets local resistance — AccBalanced · 2026-08-29
- Running Qwen 3.8 MoE on a single DGX Spark: A practical recipe — QuixiAI · 2026-08-29
- Qwen 3.8 Flash NVFP4 deployment config tested on single DGX — QuixiAI · 2026-08-29
- AMD releases ROCm 10.0, built for the Age of Agentic AI — pmttyji · 2026-08-29
- Kyndryl and Broadcom bet on private AI clouds, with certified talent as the key — DavidLinthicum · 2026-08-29