vLLM Day-0 Support for Qwen3.8-2.4T-A95B: Runs on Single NVIDIA/AMD Nodes
vllm_project · x · 2026-08-13
Congratulations to Alibaba Qwen on the release of Qwen3.8-2.4T-A95B, one of the largest open-weight models to date, featuring 2.4T total parameters, 95B active parameters, and 512 experts.
vLLM project has announced Day-0 support for this massive model, verified on both NVIDIA and AMD hardware. Ready-made 4-bit checkpoints are provided out of the box:
- NVIDIA: Runs on a single 8xB300 node (1.32 TiB).
- AMD: Runs on a single 8xMI355X node (1.45 TiB).
Related event: vLLM Announces Day-0 Support for Qwen3.8 2.4T Model(2 posts)→
More from Infra
- SkyPilot Unifies Multiple Slurm Clusters, Solving GPU Management Bottlenecks — skypilot_org · 2026-08-13
- 2-bit Quantized Nemotron 3.5 Runs Autonomous Tool Calls Continuously on Just 22GB VRAM — danielhanchen · 2026-08-13
- Report: SpaceXAI Builds Custom GB300 Inference Stack for 2× Performance Gains — XFreeze · 2026-08-13
- Tracking Amazon Bedrock Costs with Athena and CUDOS Dashboards — AWS ML Blog · 2026-08-13
- Running Minimax H3 Locally: A 6GB RTX 3050 VRAM Test — JadedScorpion · 2026-08-13
- AI Boom Spreads: Investors Target Chip Fab and Data Center Suppliers — Polymarket · 2026-08-13