Running Qwen3.8-Flash-Next on a 96GB Mac Studio: A Deep Dive

Mxmtm · reddit · 2026-08-31

A detailed technical analysis of fitting the Qwen3.8-Flash-Next model onto a 96GB Mac Studio (M3 Ultra). The post breaks down memory math for the MoE architecture, the n-gram embedding table, and KV Cache. It compares various quantization builds (Unsloth, AtomicChat, MLX) and raises critical questions about Metal's mmap behavior, tensor splitting, and framework choice between llama.cpp and MLX.

Original post →

More from Infra

Infra channel →