M3 Max Runs 125B Qwen Model Locally at 70 tok/s

A developer ran the Qwen 3.8 Flash Next (125B-A6B) model natively via mlx-serve on an M3 Max with 128GB, achieving 70 tok/s inference and 300 tok/s prefill, with 20GB memory still free after 32k-token tasks despite thermal throttling.

2026-08-28 ~ 2026-08-28 · 2 related posts