Running Qwen3.8-27B on an RTX 5090 with vLLM: NVFP4 works, FP8 doesn't
BitGreen1270 · reddit · 2026-10-11
A Redditor shares a working vLLM setup for Qwen3.8-27B on an RTX 5090: the FP8 checkpoint (30.9GB) is impractical due to VRAM and minimum-context limits, but the QUASAR-QAT NVFP4 build runs great with images or up to 256k context. The post includes a complete config: expandablesegments, 8GB KV cache cap, fp8 KV cache, prefix caching, qwen3coder tool parser, and MTP speculative decoding with 3 tokens.
More from Infra
- Dev in prod: running a newsletter side project on a cheap persistent VM — davidcrawshaw · 2026-10-11
- Zeeg: persistent VMs for agents are wrong, ephemeral sandboxes are the present — zeeg · 2026-10-11
- OpenRouter Processes ~$1.5B Annualized Inference, Anthropic Takes 36% of Spend — deedydas · 2026-10-11
- AMD Reportedly Raises GDDR6 Prices for Board Partners Amid GPU Price Pressure — chemist_slime · 2026-10-11
- M5 Ultra Mac Studio: is 80-core worth +$1200 and 10 extra weeks over 64-core? — Porespellar · 2026-10-11
- Virginia animal shelter says nearby data center noise is terrifying its rescue dogs — Polymarket · 2026-10-11