Qwen3.8-Next-Flash hits ~240 t/s decode on a single RTX 6000 Pro

AdventurousSwim1312 · reddit · 2026-08-31

Building on jpezzulli's patched sglang (178 t/s), a Reddit user let an AI iterate for days and pushed Qwen3.8-Flash-Next-NVFP4 decode to nearly 240 t/s on a single RTX 6000 Pro MaxQ (300W), vs a theoretical nvfp4 bandwidth limit around 280 t/s.

Tricks used:

Next up: full nvfp4 layer quantization, layer-fusion kernels, MTP post-training and an eagle3 head. Repro resources (patch repo, base repo, model checkpoint) are all open. Test rig: Ryzen 9 3950X, 128GB RAM (ngram table), 1x RTX 6000 Pro MaxQ.

Original post →

More from Infra

Infra channel →