DIY vLLM Recipe Runs Qwen38 Flash Next NVFP4 on RTX Pro 5000 72GB

knob-0u812 · reddit · 2026-09-26

Unable to find a ready recipe, the author combined Hermes and Unsloth's 4-bit quants to build a vLLM setup for Qwen38 Flash Next NVFP4 (TP=2, PLE offloading) on an RTX Pro 5000 72GB, running stably for a week. Config is open-sourced on GitHub.

Related event: Dev Shares vLLM Recipe for Qwen3.8 Flash Next NVFP4 on RTX Pro 5000(2 posts)→

Original post →

More from Infra

Infra channel →