NVIDIA demos 2,529 output tokens per second for Qwen 27B at GTC

IanAndrewsDC · x · 2026-09-28

NVIDIA's Ian Buck took the stage at a GTC event, reporting inference performance of 2,529 output tokens per second for Qwen 3.8 27B. The figure drew attention from folks at SemiAnalysis, Cerebras, and Inferact, offering a benchmark point for comparing inference-stack throughput across providers.

Original post →

More from Infra

Infra channel →