180B-Class Qwen3.8 Model Runs on Just 39GB Memory
EyalToledano · x · 2026-08-28
From the HamsterResearch lab: Qwen3.8-Flash-Next-REAP-288-MLX-4bit is a 180B-class model running on just 39GB of memory.
- MLX-native 4-bit, 60% smaller than stock q4.
- Pruned 512→288 experts via REAP.
- 91.5% HumanEval (vs 93.9% stock).
Related event: Pruning Cuts 180B MoE Model to Run in 39GB on Mac(2 posts)→
More from Infra
- OpenAI's Jalapeño inference chip: 9 months to tapeout, one chip beats GPU+LPU pair — thehiphopswami · 2026-08-28
- Intel XE3P projected specs: 1.3 PFLOPS FP8, 1.5TB/s bandwidth, 2027 launch — QuixiAI · 2026-08-28
- Rumor: Anthropic interested in developing its own training chip — zephyr_z9 · 2026-08-28
- India Commits $13.4B for 'Semicon 2.0' Chip Design and Manufacturing — SumitGup · 2026-08-28
- Meta, Google, NTT to discuss AI data center optical architectures — jwt0625 · 2026-08-28
- Coherent-lite Catches Up to IMDD in Energy Efficiency for Pluggables — jwt0625 · 2026-08-28