Running 262K context on a 5090+4070 rig: three Strata patches hit 130 tok/s decode
Fz1zz · reddit · 2026-10-03
A Reddit user details running Qwen3.8-Flash-Next (ISTA GSQ-RCO IQ3XXS) at full 262K context on a RTX 5090 + RTX 4070 Ti SUPER (chipset PCIe x1, 0.8 GB/s) with only 31GB RAM, using the Strata engine to stream MoE experts from an mmap'd GGUF.
- Baseline: stock v0.1.33 with a 36-layer split gave 890 tok/s on 80K prompts; switching between two chats took 20-48s
- Three self-written patches:
- All prompt prefill on the big card (port of Strata PR #269), then copy only used cells to the small card — conversation parking drops to 0.6-1.2s
- Under memory pressure MADVWILLNEED readahead fails, causing millions of major page faults; per-slice pread cuts faults per benchmark run from 24M to 14K
- Bigger prefill chunks (--prefill auto:32768) cut chunk count 4x, avoiding repeated expert re-streaming
Final config (int8 KV, 32K cells/layer, MTP spec 4, vision on): 80K prompt at 1796 tok/s (45s), median decode 128-134 tok/s, follow-up turn on an 80K conversation in 6s; 40K needle and two-conversation parking tests pass. Patches, install script and benchmarks are open-sourced: github.com/ExTV/strata-5090-4070.
More from Infra
- Suhail hails the start of the Vera Rubin era as NVIDIA's next-gen GPU boots up — rickasaurus · 2026-10-03
- GPU-backed loans put GPU earning power and collateral value under scrutiny — AnneliesGamble · 2026-10-03
- Google accused of illegally bulldozing 300 million sq. meters of Finnish forest for AI data centers — Polymarket · 2026-10-03
- SkyRL v0.4 trains 1T-param Kimi K2.7 with RL on just 16 B300 GPUs — casper_hansen_ · 2026-10-03
- BIS report: 55% of AI investment is circular, echoing Lucent-Nortel era risks — rohanpaul_ai · 2026-10-03
- RAM prices may 10-20x next year and double again by 2028, predicts analyst — Kyrannio · 2026-10-03