Building a Multi-GPU Workstation for Local 122B LLM Inference on a Budget
whatyathinkk · reddit · 2026-08-14
A developer sought advice on buying a refurbished Lenovo P620 workstation (with a Threadripper Pro 3975WX) for €700, planning to add two 5060Ti GPUs to run and serve LLMs locally.
- Current Bottleneck: Aims to run a 122B parameter model headlessly, but the current rig crashes due to VRAM and RAM limits.
- Upgrade Plan: Wants to add two 16GB GPUs for concurrent inference, though unsure if system RAM needs to scale proportionally.
- Future Goal: Hopes to incrementally buy components to eventually run massive models locally.
Replies focused on VRAM sharing mechanisms across multiple GPUs, PCIe lane bottlenecks, and the performance hit from slower DDR4 memory.
More from Infra
- Hugging Face's datatrove 0.10.0 adds HF Jobs pipeline executor, no Slurm needed — vanstriendaniel · 2026-08-14
- Hosting a 3T Model on Cerebras Requires 68 Chips and 1.5MW — zephyr_z9 · 2026-08-14
- $1T of AI Infra Buildout to Solve $1M Math Problems? — suchenzang · 2026-08-14
- Intel and MiniMax Release MXFP4 Quantized Model, Topping FP4 Benchmarks — HaihaoShen · 2026-08-14
- GPU Prices Surge Again: RTX 6000 Pro Jumps to $15k — HankYeomans · 2026-08-14
- No Real GPU Shortage, But Rental Terms Are Brutal: 3-5 Year Commits — AccBalanced · 2026-08-14