Pruning Qwen to 288 experts enables 180B model on Mac with 39GB RAM
EyalToledano · x · 2026-08-28
The author experimented with structural pruning on Qwen3.8-Flash-Next (180B MoE). By analyzing expert activity in real coding sessions, experts per layer were reduced from 512 to 288. Combined with MLX 4-bit quantization, this maintains 91.5% HumanEval accuracy while dropping memory requirements to 39GB (via streaming n-gram table), enabling the model to run on Apple Silicon.
Related event: Dev prunes and quantizes 180B-class Qwen3.8 to run on a MacBook(8 posts)→
More from Infra
- Australia Minister: No Fossil Fuel Carve-out for Datacenters — nordicinst · 2026-08-28
- ZED Camera priced at $500? DIY alternative costs just $150 — _William_F_ · 2026-08-28
- KOTOR Remaster Path Tracer Integrates DLSS 4.5 RR — Michael_Moroz_ · 2026-08-28
- NVIDIA's NVHBM Breaks the Die-Size Limit, 3-5x VRAM per GPU — Charuru · 2026-08-28
- Tutorial: Train a Raspberry Pi to Read Gas Meter Automatically with Neural Network — JeremyCMorgan · 2026-08-28
- Ninfer Benchmark: 5090 Doubles Throughput for Qwen3 27B — Rollingsound514 · 2026-08-28