Can 96GB Mac Studio run Qwen3.8? Analyzing SSD offload feasibility

Mxmtm · reddit · 2026-08-27

A user plans to buy a 96GB Mac Studio to run the 125B Qwen3.8-Flash-Next, relying heavily on llama.cpp's PLE/n-gram SSD offload feature. The post details memory calculations (Model 103.8 GiB + KV Cache 8 GiB), concluding that without offload, 128GB RAM is required, but with it, the 96GB model (84 GiB usable) might suffice. The user discusses the current implementation status (uncertain Metal I/O support), limitations of MLX, and refutes pessimistic claims of 1GB/token traffic, estimating actual load at 256KB/token.

Original post →

More from Infra

Infra channel →