Padding trick lets vLLM run tp=6: 27B model at 50 tok/s on six 7900 XTX GPUs
Biomass23 · reddit · 2026-10-01
A user shares a practical vLLM trick: tensor parallel normally requires model dimensions divisible by GPU count (power-of-two setups). By writing a converter that zero-pads the model until all required numbers divide by 6, they run a Qwen 27B model at BF16 with 256k context across six 7900 XTX GPUs.
Results:
- 50 tok/s single user, 200 tok/s aggregate with 8 concurrent prompts
- GPU KV cache of 520,784 tokens; 1.99x concurrency for 262,144-token requests
A directly reproducible approach for non-power-of-two GPU counts wanting maximum KV cache.
More from Infra
- Paraguay's $12B Iguazú AI City to host Latin America's largest data center — teortaxesTex · 2026-10-01
- You.com, NVIDIA and CoreWeave bring live web search into RL training, starting with Nemotron 3.5 Lightning — RichardSocher · 2026-10-01
- Chutes team on Bittensor Subnet 64 may have stumbled on a new way to train models while fixing inference economics — markjeffrey · 2026-10-01
- Ornith-1.5 DFlash draft models deliver up to 2.54x lossless inference speedup — alan_ritter · 2026-10-01
- Manager locked Teams transcripts, employee used Copilot to dig JSON URL out of page source — TheBigCrowbroski · 2026-10-01
- Compute per MW comparison: Nvidia still best price/perf despite prices — Storge2 · 2026-10-01