Running Qwen-Image-2.1 on a 48GB Mac: staged loading cuts peak memory from 43.6GB to 19GB

nefayran · reddit · 2026-10-07

Running Qwen-Image-2.1 on a 48GB M5 Pro Mac, a developer hit memory blowups: the default diffusers pipeline keeps all three models resident (16.3GB Qwen3-VL text encoder, 13.3GB transformer, 1.3GB VAE), pushing peak usage to 43.6GB with 8GB of swap before crashing.

Since Mac CPU and GPU share the same RAM, offloading idle models to CPU doesn't lower the peak. The fix: stageload, a small MIT-licensed library that hooks encodeprompt, preparelatents and unpacklatents to load only each stage's models and free the rest — no pipeline code changes needed.

Results at 1024×1024, 20 steps, bfloat16:

Install with pip install stageload. The approach applies to any multi-model Python pipeline on Mac.

Original post →

More from Infra

Infra channel →