Running Qwen-Image-2.1 on a 48GB Mac: staged loading cuts peak memory from 43.6GB to 19GB
nefayran · reddit · 2026-10-07
Running Qwen-Image-2.1 on a 48GB M5 Pro Mac, a developer hit memory blowups: the default diffusers pipeline keeps all three models resident (16.3GB Qwen3-VL text encoder, 13.3GB transformer, 1.3GB VAE), pushing peak usage to 43.6GB with 8GB of swap before crashing.
Since Mac CPU and GPU share the same RAM, offloading idle models to CPU doesn't lower the peak. The fix: stageload, a small MIT-licensed library that hooks encodeprompt, preparelatents and unpacklatents to load only each stage's models and free the rest — no pipeline code changes needed.
Results at 1024×1024, 20 steps, bfloat16:
- Peak memory down to 19.0GB, zero swap; decode needs only 14.4GB
- Outputs bit-for-bit identical to eager across two staged runs
- Denoising at 79s vs eager's 77.5s — nearly free
- In-stage loading (13s) beats eager's upfront 20s
Install with pip install stageload. The approach applies to any multi-model Python pipeline on Mac.
More from Infra
- Datology AI open-sources Zephon, a deterministic on-the-fly dataloader that fixes cross-GPU-run inconsistency — josh_wills · 2026-10-07
- Project Maya runs 321B GLM-5.3-Flash locally on two old V100s at up to 40 tok/s — lxfater · 2026-10-07
- Data center water and power fears overblown? Aluminum smelter uses a Boston-sized grid — csuwildcat · 2026-10-07
- Omarchy Linux ships official browser-based remote desktop plugin — juntao · 2026-10-07
- Intel CEO confirms continued chip partnership with Musk's Terafab project — elonmusk · 2026-10-07
- Symmetrix-XL: open-source engine simulates 10M atoms on a single GPU — CatAstro_Piyush · 2026-10-07