FULL STORY
Pruning Qwen3's 180B Giant Down to Laptop Size
Developers pruned Qwen3.8-Flash-Next (a 180B-class MoE) to run on consumer hardware: REAP pruning cut 25% of experts, then Eyal Toledano got it running on a 48GB MacBook with Q4 quantization.
2026-08-27 ~ 2026-08-28 · 2 episodes · 4 posts
Episode 1 · Pruned Qwen3 180B MoE Runs on Laptop with 65GB Footprint (2026-08-27, 2 posts)
Developers pruned the 180B-class Qwen3.8-Flash-Next MoE using REAP, cutting memory from 97GB to 65GB, and combined it with NVMe offload to run the model on a single 48GB GPU or even a laptop.
- Pruning Qwen3 MoE to 65GB Fits a 180B-Class Model on a 128GB Laptop — EyalToledano · 2026-08-27
- Running 180B Models on 48GB VRAM via NVMe Offload and Expert Pruning — EyalToledano · 2026-08-27
Episode 2 · Developer Fits Qwen3.8-Flash-Next into 48GB MacBook (2026-08-27, 2 posts)
Developer Eyal Toledano used pruning and lookup-table techniques to run Qwen3.8-Flash-Next in 39GB of memory via Q4 quantization on MLX, later clarifying that the 28 tok/s speed applies to high-bandwidth machines like the M4 Max, not the MacBook Air.
- Qwen3.8-Flash-Next Fits on a 48GB MacBook: Pruning + SSD N-gram Table, 39GB RAM — EyalToledano · 2026-08-27
- Two tricks squeeze Qwen3.8-Flash-Next onto a 48GB MacBook Air — EyalToledano · 2026-08-28