Qwen3.8-Next-Flash hits ~240 t/s decode on a single RTX 6000 Pro
AdventurousSwim1312 · reddit · 2026-08-31
Building on jpezzulli's patched sglang (178 t/s), a Reddit user let an AI iterate for days and pushed Qwen3.8-Flash-Next-NVFP4 decode to nearly 240 t/s on a single RTX 6000 Pro MaxQ (300W), vs a theoretical nvfp4 bandwidth limit around 280 t/s.
Tricks used:
- Further quantizing lm head and several layers from bf16 to fp8 to cut bandwidth
- Kernel tuning of selected layers for GPU efficiency
- MTP config tuning
Next up: full nvfp4 layer quantization, layer-fusion kernels, MTP post-training and an eagle3 head. Repro resources (patch repo, base repo, model checkpoint) are all open. Test rig: Ryzen 9 3950X, 128GB RAM (ngram table), 1x RTX 6000 Pro MaxQ.
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01
- Data Center Worker: Fastest Blue-Collar Path to Six Figures Right Now — AICopyLab · 2026-09-01