5 dual-L40 hosts running vLLM leave ~5GB idle VRAM per GPU — how to use it?

gulensah · reddit · 2026-08-18

A Reddit ops discussion: the user manages 5 physical hosts (two NVIDIA L40s each) with Proxmox, GPU passthrough to Ubuntu VMs, and multiple vLLM instances in Docker on each card. VRAM utilization sits around 90%, leaving 5GB free per card. Since model sizes vary across cards, they can't achieve 100% packing and ask whether there's an elegant way to use these fragmented VRAM slices across cards.

Original post →

More from Infra

Infra channel →