vLLM recipe for Qwen3.8 Flash Next NVFP4 with TP=2 on RTX Pro 5000 72GB

knob-0u812 · reddit · 2026-09-26

The author couldn't find a deployment recipe for Qwen3.8 Flash Next on an RTX Pro 5000 72GB, so they built one: an NVFP4 quant with PLE offloading in vLLM, derived from Hermes and Unsloth's 4-bit quants of the same model.

They've been running the model with this config for about a week, handling everything thrown at it, and are happy with the performance. The recipe is published in the GitHub repo qwen38-flashnext-nvfp4-sm120, with feedback welcome — useful for anyone deploying the quantized model locally on 72GB cards.

Related event: Dev Shares vLLM Recipe for Qwen3.8 Flash Next NVFP4 on RTX Pro 5000(2 posts)→

Original post →

More from Infra

Infra channel →