NVIDIA demos 2,529 output tokens per second for Qwen 27B at GTC
IanAndrewsDC · x · 2026-09-28
NVIDIA's Ian Buck took the stage at a GTC event, reporting inference performance of 2,529 output tokens per second for Qwen 3.8 27B. The figure drew attention from folks at SemiAnalysis, Cerebras, and Inferact, offering a benchmark point for comparing inference-stack throughput across providers.
More from Infra
- Well-known compute broker joins Compute Exchange as senior deals lead — ns123abc · 2026-09-29
- Chained hardcoded API key and pickle RCE gave root and full cloud takeover on a Meta service — evilsocket · 2026-09-29
- Google Cloud GA's Memorystore for Valkey 9.1 With 3x the QPS of Its Managed Redis — rseroter · 2026-09-29
- Framework opens pre-orders for 192GB Desktop with AMD Ryzen AI Max+ Pro 495 this Wednesday — gnukeith · 2026-09-29
- AWS Tutorial: Stream Qwen3-TTS Speech on SageMaker via vLLM-Omni Bidirectional Streaming — AWS ML Blog · 2026-09-29
- OriginTrail ships DKG V10.0.19 on mainnet for faster AI agent context graphs — melnykowycz · 2026-09-29