Local AI rig wisdom: 8+1 GPU layout keeps vLLM on powers of two, spare card on aux

TheZachMueller · x · 2026-09-28

gospaceport shares local AI rig advice: 8+1, 4+1 and 2+1 card layouts work well — run vLLM on the power-of-two count of GPUs, use the extra card for image/video/aux workloads, and still fit big GGUF models across all of them. TheZachMueller says this explains why odd card counts make sense, and is splitting his 9-card rig into two layers — workstations on MCIO up top, max-q cards alternating below for airflow — enabling an 8+1 setup with a 6000 ADA added.

Related event: Local Multi-GPU Builds: Pairing vLLM with Power-of-Two GPU Counts(2 posts)→

Original post →

More from Infra

Infra channel →