Single RTX PRO 4500 32GB + 64GB RAM runs 188k context at 60-67 tok/s with NVFP4
sdfprwggv · reddit · 2026-10-06
A Reddit user details running Qwen3.8-Flash-Next NVFP4 on a single RTX PRO 4500 (32GB) with just 64GB DDR5, using a Strata NVFP4 fork.
Stack highlights
- NVFP4 routed experts (63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding, W4A8 prefill on Blackwell
- Served via Strata's OpenAI-compatible server, full engine flags included
Results
- 50k context: up to 80 tok/s
- 188k warm context: 60–67 tok/s
- Cold 189k full prompt: 1,680 tok/s prefill, 53 tok/s decode
A complete, reproducible recipe for serving huge-context MoE models on consumer hardware.
More from Infra
- AI-generated key-value stores beat general-purpose DBs via specialized architecture — CShorten30 · 2026-10-06
- Only 30-40k robots installed in the US last year — the case for self-replicating factories — ihorbeaver · 2026-10-06
- Reflection says Beam hit ~19% MFU with 92% goodput during its 4-week RL run — alexpolozov · 2026-10-06
- SGLang lands layer_boundary for cross-model reuse of TP/DP/CP semantics — BanghuaZ · 2026-10-06
- Scam alert: SpotGPUs.com fakes 'transaction errors' to steal crypto deposits — Equivalent_West7788 · 2026-10-06
- How a ChatGPT-like system actually works: request-to-stream architecture explained — jawadhamza · 2026-10-06