Can a 128GB M5 Max Mac Studio Handle Concurrent Local LLM Agents?
Simple_Telephone_867 · reddit · 2026-10-06
A Reddit user awaiting an M5 Max Mac Studio (18C CPU/40C GPU/128GB) wants real-world numbers for this exact config as a local agent workstation: which 20B-70B models to run, token speeds, throughput with 3-5 agents hitting one 27B/32B model concurrently, multi-model loading/swap latency, throttling over 6-12 hour runs, and whether the real bottleneck is memory bandwidth, KV cache or compute. He plans to publish his own concurrency benchmarks on arrival.
More from Infra
- MiniMax discloses 70-80% inference margins, fueling debate on AI subscription subsidies — menhguin · 2026-10-06
- ~90% of frontier lab compute now goes to post-training and inference — IanAndrewsDC · 2026-10-06
- KLIF open-sources a Rust front-end that manages llama.cpp, vLLM and TTS servers in one window — Koksny · 2026-10-06
- Qwen 27B on 2× RX 7900 XT: 66.5 TPS single-stream, still short of claimed 100+ — EqualCryptographer67 · 2026-10-06
- Dev burns 842B tokens in September — $409k at API list price, pays just 3.4% via subscription — doodlestein · 2026-10-06
- One Dot burns 1.6B tokens/day on a $100 subscription — roughly $540k/month in API-equivalent compute — DarthSilent · 2026-10-06