How to host Qwen3.8-27b on a single RTX 5090: quants, context, and trade-offs

DustNearby2848 · reddit · 2026-08-18

A Reddit user asks how others are hosting Qwen3.8-27b for coding on a single RTX 5090. The model's extensive reasoning chains demand a lot of context; they've tried Q4 via unsloth and LM Studio, Q4 and NVFP4 via ninfer, and NVFP4 via sglang. ninfer gives the most context and speed but relies on an INT8 KV cache, which isn't ideal. They're currently experimenting with medium context and soliciting others' setups — framework, quant, and context size.

Original post →

More from Infra

Infra channel →