One RTX 5090 runs 512k context at 137 tok/s as Infernix beats Strata by 23%
zipzak · reddit · 2026-10-10
A Reddit user A/B tested local inference engines Strata and Infernix on the same model (Qwen3.8-Flash-Next Uncensored) under Windows 11 with a single RTX 5090 (32GB) and 189.6GB RAM — same-day interleaved benches, 3 runs per cell, identical inference contract (int8 KV, MTP spec-4, YaRN 2 to 512k).
Results (decode tok/s / cold prefill tok/s): at 262k, Strata Q4KS hit 120.6/1,550 vs Infernix NVFP4's 149.0/3,548; at 512k, 126.3/3,045 vs 137.0/3,497.
Takeaways:
- Infernix is +23% faster at 262k, +9% at 512k; 512k context costs ≤8% — experts, not KV, are the bottleneck on a 32GB card
- Prefix cache makes repeat long prompts nearly free (0.09s warm re-prefill of a 39k prompt)
- Quality statistically tied: PPL 4.772 vs 4.844, ΔNLL −0.015±0.010
- Vision works on both; cold start 28s (Infernix) vs 1min (Strata)
Verdict: Infernix NVFP4 (76.3GB artifact + 52GB shared n-gram volume) becomes the daily driver — 150 tok/s with 512k context and vision on one 5090. Caveats: decode measured at 39k effective context, ±10-15 tok/s noise, YaRN past 262k is experimental.
More from Infra
- $500 ex-mining BC-250 cluster runs Qwen 35B at 145 tok/s with 256k context — Ok-Breadfruit-3523 · 2026-10-11
- SpaceX reportedly starting Terafab in December: a 100M sq ft chip megafactory — bennash · 2026-10-11
- OpenAI's Jalapeño chip trades kernel difficulty for memory bandwidth, betting on AI-written kernels — nrehiew_ · 2026-10-11
- Google Web AI lead Jason Mayes grew client-side JavaScript AI usage 2500x in 5 years — jason_mayes · 2026-10-11
- Dev in prod: running a newsletter side project on a cheap persistent VM — davidcrawshaw · 2026-10-11
- Zeeg: persistent VMs for agents are wrong, ephemeral sandboxes are the present — zeeg · 2026-10-11