8+1, 4+1, 2+1: A practical GPU layout for local AI rigs with vLLM

gospaceport · x · 2026-09-28

A hobbyist shares GPU allocation tips for local AI rigs: use 8+1, 4+1, 2+1 splits, run vLLM on power-of-2 card counts, and dedicate the extra GPU to aux and image/video generation, with bonus large GGUFs stored across machines. The quoted post shows a 9-GPU rig rebuild, with the top frame sagging slightly under the weight.

Related event: Local Multi-GPU Builds: Pairing vLLM with Power-of-Two GPU Counts(2 posts)→

Original post →

More from Infra

Infra channel →