Two tricks squeeze Qwen3.8-Flash-Next onto a 48GB MacBook Air
EyalToledano · x · 2026-08-28
Correction from the author: the 28 tok/s figure applies to his M4 Max, which has enough bandwidth; a MacBook Air will decode much slower, though the model still fits within the VRAM budget.
The core setup runs q4 on MLX using only 39GB of memory (q8 needs 50GB), stacking two tricks: REAP pruning cut the expert pool from 68GB to 35GB, and the 51B n-gram table now lives on SSD instead of RAM — his patch memmaps the shards and reads 100 bytes per lookup straight off disk before dequantizing.
Related event: Developer Fits Qwen3.8-Flash-Next into 48GB MacBook(2 posts)→
More from Infra
- NVIDIA Vera CPU ships at scale to accelerate agentic workloads — rohanpaul_ai · 2026-08-28
- llama.cpp Fork Adds Qwen3.8-Flash Support, Hits 44 tks on RTX 5090 at Q4 — giveen · 2026-08-28
- antirez ports GLM 5.2 Flash to run on M5 Max, TP across two Macs — antirez · 2026-08-28
- Reddit Idea: Can a Distributed Swarm of 10x7B Coding Models Beat One 70B? — According-Extent6016 · 2026-08-28
- Lucky Robots Offers Robot Simulation Building with $50K Engineering Support — Sentdex · 2026-08-28
- Analyst: Semi Demand Exceeds Capacity by 15-20%, NVDA Supply Constrained — BenBajarin · 2026-08-28