Strata inference backend nearly doubles local decode speed vs vLLM on RTX PRO 6000
hyudryu · reddit · 2026-10-10
- A Redditor benchmarked Unsloth's Qwen3.8-Flash-Next UD-Q4KXL on a single RTX PRO 6000 (96GB) with the Strata inference backend.
- Setup: INT8 KV cache, 256K context, MTP-4 speculative decoding, 256 output tokens per task, two runs per task.
- Single-request decode speeds (tokens/s): prose 185.8, counting 320.4, coding 298.9, reasoning 289.1 — about 46-49% faster than vLLM (NVFP4) on the same card and 77-82% faster than vLLM on DGX Spark TP2.
- Author notes the quantization schemes differ so it's not perfectly apples-to-apples, but the speedup is striking; Qwen 3.8 Flash Next flies through coding tasks.
More from Infra
- One x16 slot drives 8 GPUs: blogger to benchmark PCIe switch expansion board P2P topologies — TheZachMueller · 2026-10-10
- SQLite core team builds Vec1: native ANN vector search is feature-complete — solyarisoftware · 2026-10-10
- One x16 slot, eight GPUs: how PCIe-switch expansion boards use direct-connect vs cascade topologies — TheZachMueller · 2026-10-10
- SpaceX's four AI compute deals alone add up to $41B annualized revenue — XFreeze · 2026-10-10
- Rent GPUs, self-serve 90% of tokens: dev slams unbounded LLM API pricing — TheZachMueller · 2026-10-10
- Our AI bill hit $11,400 with no attribution: a cautionary tale of unbounded retries and prompt bloat — vigilAPI · 2026-10-10