vLLM recipe for Qwen3.8 Flash Next NVFP4 with TP=2 on RTX Pro 5000 72GB
knob-0u812 · reddit · 2026-09-26
The author couldn't find a deployment recipe for Qwen3.8 Flash Next on an RTX Pro 5000 72GB, so they built one: an NVFP4 quant with PLE offloading in vLLM, derived from Hermes and Unsloth's 4-bit quants of the same model.
They've been running the model with this config for about a week, handling everything thrown at it, and are happy with the performance. The recipe is published in the GitHub repo qwen38-flashnext-nvfp4-sm120, with feedback welcome — useful for anyone deploying the quantized model locally on 72GB cards.
Related event: Dev Shares vLLM Recipe for Qwen3.8 Flash Next NVFP4 on RTX Pro 5000(2 posts)→
More from Infra
- OpenRouter launches typesafe/jev-router, a cache-aware router that picks models per request — alexcovo_eth · 2026-09-26
- Akamai beats neo-clouds on profitability, Anthropic deal and buybacks, argues investor — pdamodaran · 2026-09-26
- e/acc enthusiast nearly burns house down running his own basement GPU inference cluster — beffjezos · 2026-09-26
- SF Compute built physical settlement first, calling itself the inventor of compute markets — hardimanjames · 2026-09-26
- CUMEC's 8-inch Heterogeneous TFLN-on-SiPho: 100GHz Modulators, 1.6dB Edge Coupling Loss — jwt0625 · 2026-09-26
- Top 10% of AI Customers Drive 99.5% of Model-Serving Spend as Hyperscaler Capex Jumps $750B — rohanpaul_ai · 2026-09-26