Ask: 80GB VRAM local LLM build with three RTX 5060 Ti + Thunderbolt eGPUs

OpenEvidence9680 · reddit · 2026-08-31

A Reddit user is planning a local inference PC for llama.cpp and ComfyUI: running Qwen 3.8 Next at Q5/Q6 and GLM 3 Flash at Q4, with Darwin 31B and Qwen 27B for speed, plus dual-PC sharding and RAM offload. The plan: 3 internal RTX 5060 Ti 16GB cards (two at 8x, one at 4x) for 80GB total VRAM, 128GB RAM, and two external GPUs over Thunderbolt 5 later. He also asks whether the second PC should run Linux, and invites the community to spot flaws.

Original post →

More from Infra

Infra channel →