Mac Studio local LLM benchmarks: M5 Ultra vs Max for Qwen Flash-Next

Mxmtm · reddit · 2026-08-30

A detailed hardware comparison for running local LLMs on Mac Studio (M5 Ultra 96GB vs Max 128GB vs 64GB). The analysis covers Qwen3.8-27B and Qwen3.8-Flash-Next performance. Key findings: 64GB is non-viable for Flash-Next; 128GB Max runs Q4KXL (93.5% retention); 96GB Ultra runs IQ4XS (91.1%). The author debates bandwidth vs. VRAM trade-offs and seeks user feedback on SSD offloading, quantization quality, and GPU core impact.

Original post →

More from Infra

Infra channel →