Running 262K context on a 5090+4070 rig: three Strata patches hit 130 tok/s decode

Fz1zz · reddit · 2026-10-03

A Reddit user details running Qwen3.8-Flash-Next (ISTA GSQ-RCO IQ3XXS) at full 262K context on a RTX 5090 + RTX 4070 Ti SUPER (chipset PCIe x1, 0.8 GB/s) with only 31GB RAM, using the Strata engine to stream MoE experts from an mmap'd GGUF.

Final config (int8 KV, 32K cells/layer, MTP spec 4, vision on): 80K prompt at 1796 tok/s (45s), median decode 128-134 tok/s, follow-up turn on an 80K conversation in 6s; 40K needle and two-conversation parking tests pass. Patches, install script and benchmarks are open-sourced: github.com/ExTV/strata-5090-4070.

Original post →

More from Infra

Infra channel →