NVIDIA Details Qwen3.8-2.4T Deployment on GB300, Achieving >4K Tokens/s per GPU

PyTorch · x · 2026-08-21

NVIDIA published a technical blog detailing the deployment of Alibaba's latest open-weight model, Qwen3.8-2.4T-A95B, on the NVIDIA GB300 NVL72 platform.

Original post →

More from Infra

Infra channel →