Running 180B Models on 48GB VRAM via NVMe Offload and Expert Pruning

EyalToledano · x · 2026-08-27

Combining NVMe offload (PLE) and expert pruning (REAP) makes it possible to run 180B-class models on a single 48GB GPU. Technical details include:

This approach is particularly suitable for the MLX framework and allows further optimization by moving processes to the CPU.

Related event: Pruned Qwen3 180B MoE Runs on Laptop with 65GB Footprint(2 posts)→

Original post →

More from Infra

Infra channel →