$3,000 for local LLMs: Mac mini M5 Pro 64GB unified memory vs dual RTX 5060 Ti
AdRepulsive7837 · reddit · 2026-08-26
With a $3,000 budget for local LLM inference and agentic workloads, the user is choosing between a Mac mini M5 Pro (15-core CPU/16-core GPU, 64GB unified memory, 306GB/s bandwidth) and dual RTX 5060 Ti 16GB (32GB total VRAM, 448GB/s bandwidth). He previously used both an RTX 4090 and a Mac Studio M3 Ultra.
Mac pros: 64GB fits larger models and longer context; compact, quiet, power-efficient, cooler, and the Mac local-LLM ecosystem is improving fast. Dual-GPU pros: hands-on experience with Blackwell, CUDA, VLMs and inference frameworks like SGLang — more educational value, though 32GB VRAM feels limiting for GenAI.
Typical workload: Qwen3.8 27B to 30B models, medium-context agent work (Pi Coding agent, up to 50k tokens per project), plus casual chat. He asks the community which to pick.
More from Infra
- Local AI Registry: Open source index for hardware, models, and deployment recipes — StefanoGogioso · 2026-08-26
- TRANSIT runtime cuts LLM training GPU needs by up to 50% — PyTorch · 2026-08-26
- Ollama Announces GLM-5.3-Flash Coming Soon to Cloud Service — ollama · 2026-08-26
- Polymarket: 13% Chance AI Bubble Bursts by End of 2026 — Polymarket · 2026-08-26
- Open-Source GPU Price Aggregator Vram Watch Released — KyeGomezB · 2026-08-26
- OpenAI plans world's largest data center in Ohio, requiring more power than all state homes — bennash · 2026-08-26