Mac Studio local LLM benchmarks: M5 Ultra vs Max for Qwen Flash-Next
Mxmtm · reddit · 2026-08-30
A detailed hardware comparison for running local LLMs on Mac Studio (M5 Ultra 96GB vs Max 128GB vs 64GB). The analysis covers Qwen3.8-27B and Qwen3.8-Flash-Next performance. Key findings: 64GB is non-viable for Flash-Next; 128GB Max runs Q4KXL (93.5% retention); 96GB Ultra runs IQ4XS (91.1%). The author debates bandwidth vs. VRAM trade-offs and seeks user feedback on SSD offloading, quantization quality, and GPU core impact.
More from Infra
- 40nm Neural-Dynamics Chip Uses Conductance Drift for 2.12ms Iteration Latency — maier_ak · 2026-09-01
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01