How to host Qwen3.8-27b on a single RTX 5090: quants, context, and trade-offs
DustNearby2848 · reddit · 2026-08-18
A Reddit user asks how others are hosting Qwen3.8-27b for coding on a single RTX 5090. The model's extensive reasoning chains demand a lot of context; they've tried Q4 via unsloth and LM Studio, Q4 and NVFP4 via ninfer, and NVFP4 via sglang. ninfer gives the most context and speed but relies on an INT8 KV cache, which isn't ideal. They're currently experimenting with medium context and soliciting others' setups — framework, quant, and context size.
More from Infra
- Snapdragon X2 Elite Extreme Doubles AI Performance Over Rivals — ryanshrout · 2026-08-18
- Salvaging parts: building a local AI rig with 64GB RAM and a €1700 budget — joquinjack · 2026-08-18
- Mesh LLM: A Third Option Between Crypto Rigs and Cloud Subscriptions — alex_verem · 2026-08-18
- Overclocking VRAM on 4x RTX 5060Ti for LLM inference — Ok-Breakfast1878 · 2026-08-18
- Qwen3.8-9B MLX port runs on 16GB Macs with fast speed — alexcovo_eth · 2026-08-18
- Qwen 3.8 on Apple Silicon speeds up nearly 3x using AI-written kernels — alexcovo_eth · 2026-08-18