A working vLLM recipe runs nvfp4 KV cache on 2× RTX 5060 Ti

gtrak · reddit · 2026-07-21

A Reddit post shares a working vLLM launch recipe for running nvfp4 KV cache on 2× RTX 5060 Ti with a Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm model.

What the setup uses

Operational notes

The main value here is a concrete, reproducible serving configuration for squeezing long-context inference and speculative decoding into a consumer GPU setup.

Original post →

More from Infra

Infra channel →