Running Qwen3.8-27B on an RTX 5090 with vLLM: NVFP4 works, FP8 doesn't

BitGreen1270 · reddit · 2026-10-11

A Redditor shares a working vLLM setup for Qwen3.8-27B on an RTX 5090: the FP8 checkpoint (30.9GB) is impractical due to VRAM and minimum-context limits, but the QUASAR-QAT NVFP4 build runs great with images or up to 256k context. The post includes a complete config: expandablesegments, 8GB KV cache cap, fp8 KV cache, prefix caching, qwen3coder tool parser, and MTP speculative decoding with 3 tokens.

Original post →

More from Infra

Infra channel →