Running 104GB Qwen Model on 48GB Mac via SSD Streaming
carloslfu · hn · 2026-09-02
The author presents slotstream, a method to run the 125B parameter Qwen3.8-Flash-Next 4-bit model on Macs with as little as 16GB RAM—normally requiring 100GB+ memory. It leverages expert-offloading and SSD streaming. Built with Apple's MLX and Swift, it features an auto-mode for balancing memory and speed, with speculative decoding support planned.
More from Infra
- DGX Spark owners flag bug: latest CUDA doesn't ship the instant it's released — QuixiAI · 2026-09-03
- VideoDeltaNet open-sources hybrid attention that speeds up MiniMax H3 video generation up to 90x — realmrfakename · 2026-09-03
- Analyst: NVIDIA Could Become Intel Foundry's 'Customer Zero' as a Second Source Beyond TSMC — BenBajarin · 2026-09-03
- Investors bullish on Meta as Muse Spark 1.3 pricing undercuts frontier rivals — Scobleizer · 2026-09-03
- Fervo hits 1,064 MW under contract as Google takes option on 600 MW more — aronchick · 2026-09-03
- METR Publishes Investigation Report on OpenAI / Hugging Face Hacking Incident — stikit · 2026-09-03