Power-limited 5090 + 96GB DDR5 hits 150-200 tok/s decode on Qwen3.8-Flash-Next locally
z0_o6 · reddit · 2026-10-02
A Reddit user reports running Strata on a power-limited RTX 5090 with 96GB of DDR5-6400, achieving 150-200 tok/s decode and 5-6k tok/s prefill on Qwen3.8-Flash-Next at IQ3S quantization with a 128k-token context (8-bit KV cache). It shows that with partial offloading to system RAM, local long-context inference can remain very usable.
More from Infra
- Wish list: a Qwen4 27B with 100B+ Engram offloaded to RAM and NVMe for local users — casper_hansen_ · 2026-10-03
- State of Local AI 2026: one gaming GPU now matches the world's best model from Feb 2026 — Scobleizer · 2026-10-03
- Leak: DeepSeek V4.1 Pro ~2T params, trained sparse from scratch with no instabilities — teortaxesTex · 2026-10-03
- Stas Bekman demos his cluster-audit AI skill running on an AWS cluster — StasBekman · 2026-10-03
- Running 180B Qwen3.8-Flash-Next at 40-50 t/s on 64GB RAM + 16GB VRAM with Strata — danamir_ · 2026-10-03
- NVIDIA Spark price jumps 30% to $7,000, leaving local-AI builders hunting alternatives — Apprehensive_Side219 · 2026-10-03