DIY vLLM Recipe Runs Qwen38 Flash Next NVFP4 on RTX Pro 5000 72GB
knob-0u812 · reddit · 2026-09-26
Unable to find a ready recipe, the author combined Hermes and Unsloth's 4-bit quants to build a vLLM setup for Qwen38 Flash Next NVFP4 (TP=2, PLE offloading) on an RTX Pro 5000 72GB, running stably for a week. Config is open-sourced on GitHub.
Related event: Dev Shares vLLM Recipe for Qwen3.8 Flash Next NVFP4 on RTX Pro 5000(2 posts)→
More from Infra
- Quail: open-source AI-SQL engine hits 1B+ input tokens/min on a single H100 — sh_reya · 2026-09-26
- OpenRouter launches typesafe/jev-router, a cache-aware router that picks models per request — alexcovo_eth · 2026-09-26
- Akamai beats neo-clouds on profitability, Anthropic deal and buybacks, argues investor — pdamodaran · 2026-09-26
- e/acc enthusiast nearly burns house down running his own basement GPU inference cluster — beffjezos · 2026-09-26
- SF Compute built physical settlement first, calling itself the inventor of compute markets — hardimanjames · 2026-09-26
- CUMEC's 8-inch Heterogeneous TFLN-on-SiPho: 100GHz Modulators, 1.6dB Edge Coupling Loss — jwt0625 · 2026-09-26