Local AI rig wisdom: 8+1 GPU layout keeps vLLM on powers of two, spare card on aux
TheZachMueller · x · 2026-09-28
gospaceport shares local AI rig advice: 8+1, 4+1 and 2+1 card layouts work well — run vLLM on the power-of-two count of GPUs, use the extra card for image/video/aux workloads, and still fit big GGUF models across all of them. TheZachMueller says this explains why odd card counts make sense, and is splitting his 9-card rig into two layers — workstations on MCIO up top, max-q cards alternating below for airflow — enabling an 8+1 setup with a 6000 ADA added.
Related event: Local Multi-GPU Builds: Pairing vLLM with Power-of-Two GPU Counts(2 posts)→
More from Infra
- MLX MoE Layer Gets 1.5x Faster via Better Tile Scheduling in Grouped Matmul — awnihannun · 2026-09-28
- Cloudflare incident: skipped block zeroing leaked tenant data across 18 of 24 containers — arpit_bhayani · 2026-09-28
- On RTX 5090, Qwen 27B hits 200 TPS but Flash next only 50: what model sits between for coding? — MasterNomie · 2026-09-28
- Lumen Launches On-Demand Dedicated Internet Up to 100 Gbps at 10M US Sites — shashib · 2026-09-28
- Local AI comes in two flavors: laptop-scale for the masses vs SMB on-prem setups — TheZachMueller · 2026-09-28
- Cloudflare's agent-first Kitesurf browser adds WebMCP, passes 730k WPT subtests, runs in terminals — Cloudflare Blog · 2026-09-28