FULL STORY

Pruning Qwen3's 180B Giant Down to Laptop Size

Developers pruned Qwen3.8-Flash-Next (a 180B-class MoE) to run on consumer hardware: REAP pruning cut 25% of experts, then Eyal Toledano got it running on a 48GB MacBook with Q4 quantization.

2026-08-27 ~ 2026-08-28 · 2 episodes · 4 posts

Episode 1 · Pruned Qwen3 180B MoE Runs on Laptop with 65GB Footprint (2026-08-27, 2 posts)

Developers pruned the 180B-class Qwen3.8-Flash-Next MoE using REAP, cutting memory from 97GB to 65GB, and combined it with NVMe offload to run the model on a single 48GB GPU or even a laptop.

Episode 2 · Developer Fits Qwen3.8-Flash-Next into 48GB MacBook (2026-08-27, 2 posts)

Developer Eyal Toledano used pruning and lookup-table techniques to run Qwen3.8-Flash-Next in 39GB of memory via Q4 quantization on MLX, later clarifying that the 28 tok/s speed applies to high-bandwidth machines like the M4 Max, not the MacBook Air.