RTX 5090 runs Qwen3.8-27B at 262K context

Fz1zz · reddit · 2026-08-23

Detailed benchmark of running Qwen3.8-27B (NVFP4) on a single RTX 5090. With FP8 KV and prefix caching, the model fits a full 262K context in 32GB VRAM, achieving 77 tok/s decode at short context and 64.7 tok/s at 128K. Prefix caching provided a 22x speedup.

Original post →

More from Infra

Infra channel →