NInfer6000 hits 400 tok/s decode on RTX6000 running Qwen 3.8 Flash Next
lkarlslund · reddit · 2026-10-06
The author shared NInfer6000, a fork of NInfer optimized for Qwen 3.8 Flash Next on an RTX6000 96GB card, hitting 400 tok/s decode under ideal conditions and 13K tok/s prefill.
Key numbers: without speculative decoding, 8-bit gives 172 tok/s at 512 context (+46% over 16-bit). With MTP3 speculative decoding, 8K-context 8-bit reaches 381.5 tok/s, and adding --lm-head-draft pushes it to 401.3 tok/s. Prefill at 512 tokens is 12% slower in 8-bit, while 8K-length prefill is unchanged (13.9K tok/s).
The fork exists because original NInfer targets 5090-and-below cards and VLLM/llama.cpp fell short. It supports radixark's NVFP4 quants, the "Swift 1.5" variant (less thinking, same results), and vision. Open source on GitHub.
More from Infra
- Running 510GB DeepSeek V4.1 Flash on one DGX Spark: 113.6GB base plus 40MB domain sidecars — Physical_Toe_2499 · 2026-10-06
- Military AI veteran: 'AI-ready data' is a myth — data is the real bottleneck — ChinaTalk · 2026-10-06
- Hugging Face Kernels quickstart: load GPU-optimized kernels in one line — ariG23498 · 2026-10-06
- ZML inference already runs on Tenstorrent, Qualcomm, Intel and FuriosaAI chips — RemiCadene · 2026-10-06
- Drax datacentre would burn 4.9m tonnes of wood a year, emissions near double Gatwick flights — nordicinst · 2026-10-06
- Singapore data center operator DayOne files for US IPO after H1 revenue tripled to $512M — zephyr_z9 · 2026-10-06