Pruning Qwen to 288 experts enables 180B model on Mac with 39GB RAM

EyalToledano · x · 2026-08-28

The author experimented with structural pruning on Qwen3.8-Flash-Next (180B MoE). By analyzing expert activity in real coding sessions, experts per layer were reduced from 512 to 288. Combined with MLX 4-bit quantization, this maintains 91.5% HumanEval accuracy while dropping memory requirements to 39GB (via streaming n-gram table), enabling the model to run on Apple Silicon.

Related event: Dev prunes and quantizes 180B-class Qwen3.8 to run on a MacBook(8 posts)→

Original post →

More from Infra

Infra channel →