Serving & Monitoring Local LLMs Across Three Mixed-GPU Machines Without Duct Tape

ziyaulhuk12 · reddit · 2026-10-07

A Redditor running local inference on three machines — an AMD box (Ryzen 9 9950X + RX 7900 XTX), an NVIDIA box (i5-10400 + RTX 3050), and a MacBook Pro M3, all with LM Studio/Ollama — asks how to consolidate serving, routing, and monitoring. Pain points: no single view of loaded models/versions, token throughput, or VRAM/unified-memory pressure. They ask for multi-node serving solutions, observability beyond per-box Prometheus, and model version tracking, and offer to share their setup.

Original post →

More from Infra

Infra channel →