Tuning SGLang on a single 5090 for Qwen3.8-27B: 100 tok/s but only 82k context

ni1by2thetrue · reddit · 2026-09-15

A user deploying Qwen3.8-27B (Huihui abliterated NVFP4 multimodal build) on a single RTX 5090 with SGLang shared their full launch flags and asked for tuning help: 253,952 context length, fp8 KV cache, 0.97 static memory fraction, flashinfer attention backend, hierarchical write-through hicache, Mamba cache strategy, and NEXTN speculative decoding (3 steps, 4 draft tokens). They get 100 tok/s decode but only 82k usable context, and want more context without losing multimodal capability.

Original post →

More from Infra

Infra channel →