5 dual-L40 hosts running vLLM leave ~5GB idle VRAM per GPU — how to use it?
gulensah · reddit · 2026-08-18
A Reddit ops discussion: the user manages 5 physical hosts (two NVIDIA L40s each) with Proxmox, GPU passthrough to Ubuntu VMs, and multiple vLLM instances in Docker on each card. VRAM utilization sits around 90%, leaving 5GB free per card. Since model sizes vary across cards, they can't achieve 100% packing and ask whether there's an elegant way to use these fragmented VRAM slices across cards.
More from Infra
- DeepSeek V4 Flash Benchmarks: n_max=3 Yields 1.39× Speedup — Responsible_Pain3278 · 2026-08-18
- NVIDIA pledges up to $105B in project guarantees, raising questions about artificial growth fueled by self-funding chip sales — heypearlai · 2026-08-18
- Inference > Training: LLMs are served year-round but trained once or twice — prajdabre · 2026-08-18
- Testing MTP combined with ngram-mod for coding speed — YetAnotherAnonymoose · 2026-08-18
- Tesla's Optimus: Leveraging FSD and Auto Supply Chain for Humanoid Robots — 创业邦 · 2026-08-18
- With USB4STREAM merged into Linux 7.2, are inference runtimes adopting it? — voyager256 · 2026-08-18