M3 Max achieves 70 tok/s locally with 20GB RAM free after 32k task

mayfer · x · 2026-08-28

A user reported benchmark data for running LLMs locally on an M3 Max (128GB). The inference speed hit 70 tok/s with a prefill speed of around 300 tok/s. Despite thermal throttling issues, the setup remained usable for agents, completing a 32k token task while still retaining 20GB of free memory.

Related event: M3 Max Runs 125B Qwen Model Locally at 70 tok/s(2 posts)→

Original post →

More from coding & agent

coding & agent channel →