Qwen 27B NVFP4 on RTX 5090: 120 t/s with vision and 451K cache

t4a8945 · reddit · 2026-08-23

A detailed technical report on running Qwen 2.5 27B in NVFP4 quantization on a single RTX 5090 (400W limited). Using vLLM, the setup achieves 120 t/s with vision enabled and 451K global KV-cache. Includes extensive benchmarks from 4K to 185K context and a setup guide.

Original post →

More from Infra

Infra channel →