M3 Max Local Agent Test: 20GB RAM Free After 32K Task

mayfer · x · 2026-08-31

A user benchmarked local agent performance on an M3 Max (128GB). The results showed 70 tok/s inference speed and around 300 tok/s prefill. After completing a 32K token task, 20GB of memory remained free, indicating strong feasibility for running long-context agents locally on this hardware configuration.

Original post →

More from coding & agent

coding & agent channel →