Tuning SGLang on a single 5090 for Qwen3.8-27B: 100 tok/s but only 82k context
ni1by2thetrue · reddit · 2026-09-15
A user deploying Qwen3.8-27B (Huihui abliterated NVFP4 multimodal build) on a single RTX 5090 with SGLang shared their full launch flags and asked for tuning help: 253,952 context length, fp8 KV cache, 0.97 static memory fraction, flashinfer attention backend, hierarchical write-through hicache, Mamba cache strategy, and NEXTN speculative decoding (3 steps, 4 draft tokens). They get 100 tok/s decode but only 82k usable context, and want more context without losing multimodal capability.
More from Infra
- Before porting code to GPU, measure memory transfer time — data movement is the real bottleneck — Franc0Fernand0 · 2026-09-15
- MediaTek's Dimensity 9600 Pro: 2nm chip doubles AI compute, cuts power — jiqizhixin · 2026-09-15
- Four dev boards hooked to the internet: test AI-written firmware on real silicon via HTTPS — SelfishlyWandering · 2026-09-15
- Grouped Value Attention shrinks KV cache by reconstructing keys on demand — Vishesh Tripathi · 2026-09-15
- jinfer brings native AI inference to the JVM, matching llama.cpp on CPU with zero Python — mukel90 · 2026-09-15
- Wan 2.2 on one RTX 5090: frame count doesn't touch VRAM, but resolution drops it by 10GB — Realistic-Fennel-190 · 2026-09-15