Ask: 80GB VRAM local LLM build with three RTX 5060 Ti + Thunderbolt eGPUs
OpenEvidence9680 · reddit · 2026-08-31
A Reddit user is planning a local inference PC for llama.cpp and ComfyUI: running Qwen 3.8 Next at Q5/Q6 and GLM 3 Flash at Q4, with Darwin 31B and Qwen 27B for speed, plus dual-PC sharding and RAM offload. The plan: 3 internal RTX 5060 Ti 16GB cards (two at 8x, one at 4x) for 80GB total VRAM, 128GB RAM, and two external GPUs over Thunderbolt 5 later. He also asks whether the second PC should run Linux, and invites the community to spot flaws.
More from Infra
- Preparing for M5 Ultra 512GB: which model quants fit and perform best locally — Ok_Warning2146 · 2026-08-31
- Huawei's Kirin 2026 processor with LogicFolding architecture coming this fall — pstAsiatech · 2026-08-31
- Huawei revenue up 9.55% as it pours 25% into R&D for self-reliance — pstAsiatech · 2026-08-31
- Ask: does generating every token re-execute the full parameter set in an LLM? — MarinatedPickachu · 2026-08-31
- $1.07 for two days of prompting GLM-5.3 Flash shows new software costs — henkvaness · 2026-08-31
- Dual-GPU AI Workstation: R9700 for H3 + 3060 for Qwen Benchmark — Master-Client6682 · 2026-08-31