Serving & Monitoring Local LLMs Across Three Mixed-GPU Machines Without Duct Tape
ziyaulhuk12 · reddit · 2026-10-07
A Redditor running local inference on three machines — an AMD box (Ryzen 9 9950X + RX 7900 XTX), an NVIDIA box (i5-10400 + RTX 3050), and a MacBook Pro M3, all with LM Studio/Ollama — asks how to consolidate serving, routing, and monitoring. Pain points: no single view of loaded models/versions, token throughput, or VRAM/unified-memory pressure. They ask for multi-node serving solutions, observability beyond per-box Prometheus, and model version tracking, and offer to share their setup.
More from Infra
- NVFP4 runs a 2.4T-parameter model on a quarter of the GPUs, cutting per-token cost up to 73% — ryanshrout · 2026-10-07
- Intel exec: agentic AI spends real time waiting on business systems, end-to-end wall-clock is the metric that matters — ryanshrout · 2026-10-07
- Disaggregated inference is the future, says e/acc's Beff Jezos after panel with General Compute — beffjezos · 2026-10-07
- What Does 'Owning Your Own AI' Technically Mean? Locally Running Open Weights Debated — Imaginary_Choice_430 · 2026-10-07
- NVIDIA's LoGRA cuts LLM RL training memory by up to 45.7%, enables 27B RL on a single 8-GPU node — nvidia · 2026-10-07
- Ex-NVIDIA engineer tells the story of testing a chip with a broken memory controller — blelbach · 2026-10-07