Developer prunes and quantizes 180B Qwen3.8 to run on 48GB Mac

Developer Eyal Toledano completed an experiment fitting a 180B-class MoE model on a consumer Mac: based on Qwen3.8-Flash-Next, structural pruning plus MLX quantization produced a 4-bit version that runs in just 39GB of memory while keeping HumanEval at 91.5%. The experiment shows that expert redundancy in large models can be safely compressed using real-world usage data, charting a viable path to running extra-large models locally.

Confirmed

Why it matters

2026-08-27 ~ 2026-08-28 · 5 related posts

Primary sources