A vLLM NVFP4 bug fix and tuning notes for 4× RTX 5060 Ti servers
see_spot_ruminate · reddit · 2026-07-29
A Reddit user shares a concrete vLLM setup for running unsloth/Qwen3.6-27B-NVFP4 on 4× RTX 5060 Ti cards, including a workaround for an OOM startup bug.
- The reported fix for the vLLM #46268 NVFP4 startup OOM is to add both MAXJOBS=4 and NVCCTHREADS=4 to the systemd service, which slows startup but avoids the error.
- They also describe tuning gpumemoryutilization, MTP speculative decoding, and context limits to reach roughly 70–80 tok/s generation and 2,000+ tok/s prefill on a consumer multi-GPU rig.
- The post includes full software and hardware details: Ubuntu 26.04, CUDA 13.3, driver 595.71.05, NCCL, vLLM 0.26.0, and uneven PCIe lane allocation on the motherboard.
More from Infra
- LiveKit says Gemma 4 31B hits 192 ms to first token in voice agents — GlennCameronjr · 2026-07-30
- Optimized Qwen Image 2512: 5x Smaller, 3x Faster Inference — enrique-byteshape · 2026-07-30
- Reddit user gets about 4 tokens/s running Kimi K3 on a 2×5090 home lab — iVoider · 2026-07-30
- AI Infrastructure Stocks Cool Down: CRWV at 52-Week Low, NVDA Down 8.69% — GaryMarcus · 2026-07-30
- Two RTX 3090s still struggle to fit Qwen Image Edit alongside a 27B text model — Civil_Fee_7862 · 2026-07-30
- llama.cpp prefill leaves CPU cores and memory bandwidth surprisingly idle — Dependent_Ad948 · 2026-07-30