Pruned Qwen3 180B MoE Runs on Laptop with 65GB Footprint
Developers pruned the 180B-class Qwen3.8-Flash-Next MoE using REAP, cutting memory from 97GB to 65GB, and combined it with NVMe offload to run the model on a single 48GB GPU or even a laptop.
2026-08-27 ~ 2026-08-27 · 2 related posts
- Episode 1: Pruned Qwen3 180B MoE Runs on Laptop with 65GB Footprint(2026-08-27, 2 posts)
- Episode 2: Developer Fits Qwen3.8-Flash-Next into 48GB MacBook(2026-08-27, 2 posts)
- Pruning Qwen3 MoE to 65GB Fits a 180B-Class Model on a 128GB Laptop — EyalToledano · 2026-08-27
- Running 180B Models on 48GB VRAM via NVMe Offload and Expert Pruning — EyalToledano · 2026-08-27