M3 Max Runs 125B Qwen Model Locally at 70 tok/s
A developer ran the Qwen 3.8 Flash Next (125B-A6B) model natively via mlx-serve on an M3 Max with 128GB, achieving 70 tok/s inference and 300 tok/s prefill, with 20GB memory still free after 32k-token tasks despite thermal throttling.
2026-08-28 ~ 2026-08-28 · 2 related posts
- Dev runs 125B Qwen model locally on M3 Max at 70 tok/s via MLX — mayfer · 2026-08-28
- M3 Max achieves 70 tok/s locally with 20GB RAM free after 32k task — mayfer · 2026-08-28