NVIDIA Details Qwen3.8-2.4T Deployment on GB300, Achieving >4K Tokens/s per GPU
PyTorch · x · 2026-08-21
NVIDIA published a technical blog detailing the deployment of Alibaba's latest open-weight model, Qwen3.8-2.4T-A95B, on the NVIDIA GB300 NVL72 platform.
- Model Specs: The model features 2.4T total parameters with 95B active per token, utilizing a MoE architecture and hybrid attention mechanism, supporting up to 1M token context window.
- Performance: Without additional tuning, the model achieves over 4,000 tokens/s throughput per GPU and over 350 tokens/s per user in FP8 precision.
- Tooling: Developers can leverage PyTorch-native fine-tuning and NVIDIA NeMo AutoModel to perform SFT or memory-efficient LoRA fine-tuning directly on existing checkpoints without model conversion.
More from Infra
- Payman Connect Lets AI Agents Safely Access Bank Accounts and Wallets — minsuk_chang · 2026-08-21
- alt-p2p-http: Encrypted P2P tunnel for local LLM access — agf2007 · 2026-08-21
- INVAR: Open-Source Tool Gives Every Local LLM Inference a SHA-256 Receipt — revuprender · 2026-08-21
- SpaceX launch cadence could enable 12-50 GW of space compute — teortaxesTex · 2026-08-21
- Local AI Coding Hardware Tiers: $1k Gets You the Smartest Model — nickbaumann_ · 2026-08-21
- Cerebras officer Sean Lie files to sell $153M in shares as IPO retail buyers get crushed — firstadopter · 2026-08-21