RTX 5090 runs Qwen3.8-27B at 262K context
Fz1zz · reddit · 2026-08-23
Detailed benchmark of running Qwen3.8-27B (NVFP4) on a single RTX 5090. With FP8 KV and prefix caching, the model fits a full 262K context in 32GB VRAM, achieving 77 tok/s decode at short context and 64.7 tok/s at 128K. Prefix caching provided a 22x speedup.
More from Infra
- Windows installer guide for ComfyUI Trellis 2 on RDNA3/3.5/4 GPUs — Wake_Up_Morty · 2026-08-23
- Open Source Project pgrok: Self-Hosted ngrok Alternative via VPS — tom_doerr · 2026-08-23
- Tension between data center opposition and AI industry expansion — NathanpmYoung · 2026-08-23
- Upgrading RTX A6000 thermal paste and fan makes it usable for workloads — cephaloform · 2026-08-23
- Optimized llama.cpp fork for AMD GFX906 (Mi50, Mi60, Radeon VII) — milpster · 2026-08-23
- Nvidia AI Server Prices to Rise 15%+, GB300s Around $600k — zephyr_z9 · 2026-08-23