$3,000 for local LLMs: Mac mini M5 Pro 64GB unified memory vs dual RTX 5060 Ti

AdRepulsive7837 · reddit · 2026-08-26

With a $3,000 budget for local LLM inference and agentic workloads, the user is choosing between a Mac mini M5 Pro (15-core CPU/16-core GPU, 64GB unified memory, 306GB/s bandwidth) and dual RTX 5060 Ti 16GB (32GB total VRAM, 448GB/s bandwidth). He previously used both an RTX 4090 and a Mac Studio M3 Ultra.

Mac pros: 64GB fits larger models and longer context; compact, quiet, power-efficient, cooler, and the Mac local-LLM ecosystem is improving fast. Dual-GPU pros: hands-on experience with Blackwell, CUDA, VLMs and inference frameworks like SGLang — more educational value, though 32GB VRAM feels limiting for GenAI.

Typical workload: Qwen3.8 27B to 30B models, medium-context agent work (Pi Coding agent, up to 50k tokens per project), plus casual chat. He asks the community which to pick.

Original post →

More from Infra

Infra channel →