Qwen3.6 35B NVFP4 Hits 3,000 Tokens/s on a Single RTX PRO 6000

max_paperclips · x · 2026-08-05

A developer achieved 3,000 tokens/s generating speed using 16-parallel inference of the Qwen3.6 35B-A3B model in NVFP4 format on a single RTX PRO 6000 Workstation Edition GPU.

The demonstration utilized a custom tool called ParallelHue, which color-codes tokens to visualize the simultaneous generation process enabled by speculative decoding.

Original post →

More from Infra

Infra channel →