NInfer6000 hits 400 tok/s decode on RTX6000 running Qwen 3.8 Flash Next

lkarlslund · reddit · 2026-10-06

The author shared NInfer6000, a fork of NInfer optimized for Qwen 3.8 Flash Next on an RTX6000 96GB card, hitting 400 tok/s decode under ideal conditions and 13K tok/s prefill.

Key numbers: without speculative decoding, 8-bit gives 172 tok/s at 512 context (+46% over 16-bit). With MTP3 speculative decoding, 8K-context 8-bit reaches 381.5 tok/s, and adding --lm-head-draft pushes it to 401.3 tok/s. Prefill at 512 tokens is 12% slower in 8-bit, while 8K-length prefill is unchanged (13.9K tok/s).

The fork exists because original NInfer targets 5090-and-below cards and VLLM/llama.cpp fell short. It supports radixark's NVFP4 quants, the "Swift 1.5" variant (less thinking, same results), and vision. Open source on GitHub.

Original post →

More from Infra

Infra channel →