Run Qwen3.8-27B NVFP4 on RTX 5090: vLLM, 256K Context, 147 tok/s

AIFlow_ML · x · 2026-08-16

MiaAI-Lab releases a script to serve RadixArk/Qwen3.8-27B-NVFP4 with vLLM 0.27.x on a single RTX 5090 (32GB), with native 256K context, TurboQuant 4-bit KV cache (5.5 GiB), and MTP-3 speculative decoding, exposed as OpenAI-compatible API. Measured 160 tok/s single-stream generation, full 262,144-token context resident, and a patch fixing stock 0.27.1 garbling (0/15 vs 13/15 failures).

Original post →

More from coding & agent

coding & agent channel →