Local LLM Hardware Dilemma: M1 Max 64GB vs M3 Max 128GB vs 24GB GPU Build
Guilty-Support-584 · reddit · 2026-09-16
A Reddit user running Qwen3.6 35B A3B locally on an M2 Max 32GB hits severe bottlenecks—4-10 tks via SSD streaming and 20-60 minutes prompt processing for 40-60k tokens. They weigh a €1500 used M1 Max 64GB, €3500 M3 Max 128GB, or a DIY 24GB GPU + 128GB RAM build.
More from Infra
- smolperfbenchmark: a public leaderboard for small open models on 8GB-class devices — East-Muffin-6472 · 2026-09-16
- Ministral 3 3B on a Galaxy S21 relays chats between Gemini and Z.ai across two browsers, 10/10 runs — Mean-Standard7390 · 2026-09-16
- llama.cpp Merges hc Ops for Qwen4exp, Time to Re-benchmark Qwen Flash — jacek2023 · 2026-09-16
- Tencent open-sources BrowserSkill, letting AI agents borrow your logged-in browser tabs — AIFlow_ML · 2026-09-16
- Store Everything, Model Later: Data Lakehouse Lessons Applied to Context Infrastructure — blaizedsouza · 2026-09-16
- Lumentum: optics nears its biggest inflection in 30 years as it becomes part of the AI compute engine — pstAsiatech · 2026-09-16