One RTX 5090 runs 512k context at 137 tok/s as Infernix beats Strata by 23%

zipzak · reddit · 2026-10-10

A Reddit user A/B tested local inference engines Strata and Infernix on the same model (Qwen3.8-Flash-Next Uncensored) under Windows 11 with a single RTX 5090 (32GB) and 189.6GB RAM — same-day interleaved benches, 3 runs per cell, identical inference contract (int8 KV, MTP spec-4, YaRN 2 to 512k).

Results (decode tok/s / cold prefill tok/s): at 262k, Strata Q4KS hit 120.6/1,550 vs Infernix NVFP4's 149.0/3,548; at 512k, 126.3/3,045 vs 137.0/3,497.

Takeaways:

Verdict: Infernix NVFP4 (76.3GB artifact + 52GB shared n-gram volume) becomes the daily driver — 150 tok/s with 512k context and vision on one 5090. Caveats: decode measured at 39k effective context, ±10-15 tok/s noise, YaRN past 262k is experimental.

Original post →

More from Infra

Infra channel →