Qwen3-8 2.4T Hits 288k tokens/s on NVIDIA GB300 NVL72
RhubarbSimilar1683 · reddit · 2026-08-17
NVIDIA's official blog announces that the Qwen3-8 2.4T model achieves over 4K tokens per second per GPU and over 350 tokens per second per user on the NVIDIA GB300 NVL72 in FP8 precision on Day 0. Further optimizations like NVFP4 are expected to enhance performance.
More from Infra
- Data center boom drives up emissions as costs outweigh benefits — robleclerc · 2026-08-17
- DeepSeek v4 PRO Q2 hits 45 t/s with VRAM/RAM split on DGX Station — antirez · 2026-08-17
- Model routing should be optimized at the harness layer, not the gateway — agihouse_org · 2026-08-17
- Qwen3.8-27B Benchmarks: RPC inference tests across AMD and Nvidia GPUs — tabletuser_blogspot · 2026-08-17
- AI Spending Gap 625x: Top 1% Spend $7,500/Employee/Month vs Median $12 — rohanpaul_ai · 2026-08-17
- AWS to deploy Nvidia GB300 as primary GPU in 2026, expand Trainium shipments: TrendForce — Beth_Kindig · 2026-08-17