Will combining multiple GPUs' VRAM for local LLMs ever work out of the box?

PusheenHater · reddit · 2026-09-05

A Reddit discussion on the biggest bottleneck for local inference: VRAM. Substituting RAM hurts performance, and pooling multiple idle GPUs' VRAM still requires niche, complex setups. The poster asks whether easy out-of-the-box multi-GPU VRAM pooling (e.g., in ComfyUI) will ever become a reality.

Original post →

More from Infra

Infra channel →