Qwen3.6 35B NVFP4 Hits 3,000 Tokens/s on a Single RTX PRO 6000
max_paperclips · x · 2026-08-05
A developer achieved 3,000 tokens/s generating speed using 16-parallel inference of the Qwen3.6 35B-A3B model in NVFP4 format on a single RTX PRO 6000 Workstation Edition GPU.
The demonstration utilized a custom tool called ParallelHue, which color-codes tokens to visualize the simultaneous generation process enabled by speculative decoding.
More from Infra
- Deleting 90% of Weights: Song Han's Journey to Efficient AI & Quantization — JafarNajafov · 2026-08-05
- Cloudflare CEO Predicts Bot Traffic Will Reach 1,000x Human Traffic Within Five Years — 0xSammy · 2026-08-05
- DeepSeek 284B on 4x3090: Prefill Hits 1906 tok/s After Optimization — max_paperclips · 2026-08-05
- Running Minimax H3 on RTX 3060 Takes 1.5 Hours for a 15s Clip, Dev Seeks Optimization — SMPTHEHEDGEHOG · 2026-08-05
- Can RTX 5060Ti 16GB Run Qwen 27B? Users Discuss Local LLM Hardware — Yanzihko · 2026-08-05
- AI Profits Drive US Stocks to Record Highs: Palantir Revenue Jumps 93% — nordicinst · 2026-08-05