Power-limited 5090 + 96GB DDR5 hits 150-200 tok/s decode on Qwen3.8-Flash-Next locally

z0_o6 · reddit · 2026-10-02

A Reddit user reports running Strata on a power-limited RTX 5090 with 96GB of DDR5-6400, achieving 150-200 tok/s decode and 5-6k tok/s prefill on Qwen3.8-Flash-Next at IQ3S quantization with a 128k-token context (8-bit KV cache). It shows that with partial offloading to system RAM, local long-context inference can remain very usable.

Original post →

More from Infra

Infra channel →