Can 96GB Mac Studio run Qwen3.8? Analyzing SSD offload feasibility
Mxmtm · reddit · 2026-08-27
A user plans to buy a 96GB Mac Studio to run the 125B Qwen3.8-Flash-Next, relying heavily on llama.cpp's PLE/n-gram SSD offload feature. The post details memory calculations (Model 103.8 GiB + KV Cache 8 GiB), concluding that without offload, 128GB RAM is required, but with it, the 96GB model (84 GiB usable) might suffice. The user discusses the current implementation status (uncertain Metal I/O support), limitations of MLX, and refutes pessimistic claims of 1GB/token traffic, estimating actual load at 256KB/token.
More from Infra
- Nvidia Reportedly Agrees to Buy Hugging Face for $12.9 Billion — ferruz_noelia · 2026-08-27
- AI generates 500k words on 31 kWh, equaling 310 human hours of energy — cis_female · 2026-08-27
- Cursor User Burns 472.8B Tokens in a Month, Highlighting Cost Limits — Daniel_Farinax · 2026-08-27
- llama.cpp PR adds --n-cpu-ffn option for dense model offload — jacek2023 · 2026-08-27
- Deep Dive: AWS S3 Architecture, Rust Rewrite, and Heat Management at 280 Trillion Objects — Franc0Fernand0 · 2026-08-27
- Prefix Sliding: discarding stale reasoning tokens makes test-time scaling 3x faster — Bedrovelsen · 2026-08-27