DIY Local LLM Cluster: Quad RTX 5070 Ti + Quad 5060 Ti Nodes With vLLM FP8 Inference
Whahine · reddit · 2026-09-25
A New Zealand hobbyist on r/LocalLLaMA showed off a self-built local inference setup for data sovereignty: two self-contained nodes with 4×RTX 5060 Ti and 4×RTX 5070 Ti (16GB each) linked via PLX 88096 switches.
Key details: cards were collected piecemeal due to one-per-customer limits, switching to 5070 Ti when the price gap closed. Stack is Ubuntu 24.04, CUDA 13.2, vLLM 0.30 in Docker, with intra-node NCCL P2P working (cross-node still unsolved). The post includes the full vLLM launch command — Qwen3.8-27B-FP8, FP8 quantization, tensor-parallel 4, 64K context, 0.85 GPU memory utilization — plus llama-benchy benchmarks. Pain points included incompatible power adapters/PCI risers, cooling, and 2-4 week shipping with 15% import tax.
More from Infra
- Google's Project Suncatcher flies TPU prototype satellite on SpaceX Transporter-18 — Miles_Brundage · 2026-09-25
- YC-backed Isoquant launches GLM-5.3-Flash inference at $0.07/M with 452ms TTFT — ycombinator · 2026-09-25
- Nemotron 3 Speaker Diarization Ported to Apple Silicon via Core ML and MLX — ivan_digital · 2026-09-25
- Strangely, GPU matmuls run faster on 'predictable' data: Horace He explains — goyal__pramod · 2026-09-25
- PyTorch announces ExecuTorch Hackathon in San Francisco, Oct 17-18, with three device tracks — PyTorch · 2026-09-25
- Speculation: GPT-6 Luna/Sol efficiency lean hints at Cerebras 1000 tok/s inference economics — brandon_galang · 2026-09-25