M3 Max achieves 70 tok/s locally with 20GB RAM free after 32k task
mayfer · x · 2026-08-28
A user reported benchmark data for running LLMs locally on an M3 Max (128GB). The inference speed hit 70 tok/s with a prefill speed of around 300 tok/s. Despite thermal throttling issues, the setup remained usable for agents, completing a 32k token task while still retaining 20GB of free memory.
Related event: M3 Max Runs 125B Qwen Model Locally at 70 tok/s(2 posts)→
More from coding & agent
- Google DeepMind uses Teamwork multi-agent framework for math breakthroughs — algo_diver · 2026-08-28
- Oxlint Config v2 Released to Keep AI Agents on Track — cnakazawa · 2026-08-28
- Overengineering pitfalls: Misusing Docker and excessive constraints — burny_tech · 2026-08-28
- Photographer tests ChatGPT Desktop for photo editing: achieves perfect results, saves $1k/month — Ice2jc · 2026-08-28
- Video: Why teams are taking back control of their AI coding stack — labeveryday · 2026-08-28
- Relay Harness Unifies 204 Models into a Single Coding Agent with Prepaid Billing — realmeetjames · 2026-08-28