DIY Local LLM Cluster: Quad RTX 5070 Ti + Quad 5060 Ti Nodes With vLLM FP8 Inference

Whahine · reddit · 2026-09-25

A New Zealand hobbyist on r/LocalLLaMA showed off a self-built local inference setup for data sovereignty: two self-contained nodes with 4×RTX 5060 Ti and 4×RTX 5070 Ti (16GB each) linked via PLX 88096 switches.

Key details: cards were collected piecemeal due to one-per-customer limits, switching to 5070 Ti when the price gap closed. Stack is Ubuntu 24.04, CUDA 13.2, vLLM 0.30 in Docker, with intra-node NCCL P2P working (cross-node still unsolved). The post includes the full vLLM launch command — Qwen3.8-27B-FP8, FP8 quantization, tensor-parallel 4, 64K context, 0.85 GPU memory utilization — plus llama-benchy benchmarks. Pain points included incompatible power adapters/PCI risers, cooling, and 2-4 week shipping with 15% import tax.

Original post →

More from Infra

Infra channel →