M3 Max Local Agent Test: 20GB RAM Free After 32K Task
mayfer · x · 2026-08-31
A user benchmarked local agent performance on an M3 Max (128GB). The results showed 70 tok/s inference speed and around 300 tok/s prefill. After completing a 32K token task, 20GB of memory remained free, indicating strong feasibility for running long-context agents locally on this hardware configuration.
More from coding & agent
- Arctron AI update: Visual/accuracy improvements and WebMCP support — jasonkneen · 2026-08-31
- Shipstatic MCP server lets AI agents deploy and manage static sites — modelcontextprotocol · 2026-08-31
- New MCP connector enables LLM agents to perform infrastructure management — modelcontextprotocol · 2026-08-31
- Visual summary checks fail for LLM-edited formulas at scale; need diffs and dependency checks — saimatrixxx · 2026-08-31
- 12-layer engineering stack for production AI agents derived from 400+ builds — MaryamMiradi · 2026-08-31
- Assigning different models to specific agent roles: Luna for info, Sol/Opus for coding, Fable for planning — MilesCranmer · 2026-08-31