Pruning Qwen3 MoE to 65GB Fits a 180B-Class Model on a 128GB Laptop
EyalToledano · x · 2026-08-27
- Early results pruning Qwen3.8-Flash-Next experts with REAP look very promising:
- REAP-384 (25% experts pruned): 80GB vs. 97GB stock, 89% accuracy retained;
- REAP-256 (50% pruned): 65GB vs. 97GB, 81% accuracy, and loads twice as fast.
- Interesting tidbit: REAP-256 became chattier and did more thinking than stock.
- Bottom line: a 128GB laptop can now load a 180B-class model at 65GB with 40GB+ headroom for long context — potentially daily-driver territory.
More from Infra
- ABF Substrates & PCB Identified as Key Constraints for 2027 AI Hardware — zephyr_z9 · 2026-08-27
- First startup enriches uranium for nuclear-powered data centers — Polymarket · 2026-08-27
- G2 spent $1.27M on 970B tokens, shifting focus from adoption to efficiency — prasanna_says · 2026-08-27
- Zhipu's GLM 3.5 Flash Served 42T Tokens in 6 Days Free on Chinese Chips — bindureddy · 2026-08-27
- TokenVisor supports Nvidia, AMD, and Intel GPUs in a single cluster — AccBalanced · 2026-08-27
- Zai's domestic inference cluster hits 100k+ chips; GLM-5.3 runs on custom interconnect — zephyr_z9 · 2026-08-27