Enthusiast Shares 3-Year Journey of Building 4x RTX 6000 Pro Local AI Cluster

Tourus · reddit · 2026-08-09

A developer detailed the 3-year evolution of their local AI cluster, scaling from a single gaming GPU to 4x RTX 6000 Pro Max Q + 4x RTX 3090s. The core motivation was keeping private data offline and ensuring stable local inference.

The author shared extensive hardware troubleshooting experiences: low-quality PSUs and PCIe cables caused GPUs to drop off the bus and nearly started a fire. He notes that cloud computing is undeniably cheaper for pure compute; the local build is strictly for privacy and enthusiast purposes. Currently, 90% of the cluster's usage is for startup development, completely eliminating API rate limits and key leakage worries, while smoothly running models like GLM 5.2 and various image/voice workflows.

Original post →

More from Infra

Infra channel →