vLLM Announces Day-0 Support for Qwen3.8 2.4T with Ready-to-use 4-bit Checkpoints
vllm_project · x · 2026-08-13
The vLLM project has announced Day-0 support for Alibaba's massive new open-weight model, Qwen3.8-2.4T-A95B (2.4T total params, 95B active, 512 experts).
Verified on both NVIDIA and AMD hardware, vLLM has collaborated with Inferact to provide ready-made 4-bit quantized checkpoints:
- NVFP4 (1.32 TiB): Runs on a single 8xNVIDIA B300 node.
- MXFP4 (1.45 TiB): Runs on a single 8xAMD MI355X node.
Users can deploy directly using vllm serve without needing any manual conversion or calibration.
Related event: vLLM Announces Day-0 Support for Qwen3.8 2.4T Model(2 posts)→
More from Infra
- Wetty: Run a Terminal Emulator in Your Browser via Node.js and SSH — tom_doerr · 2026-08-13
- SK Hynix to Invest $38.1B in Two New Memory Fabs in Korea — Beth_Kindig · 2026-08-13
- Emerging Inference Engine 'Tokenspeed' Gains Day-0 Support for Qwen3.8 — This_Maintenance_834 · 2026-08-13
- DeepSeek V4 Analysis: Ultra-low API Costs and SSD KV Cache Revolution — Xianbao_QIAN · 2026-08-13
- Vercel AI Gateway Adds Grok and DeepSeek Models with Zero Markup — brandon_galang · 2026-08-13
- Rethinking AI Energy Metrics: Joules per Token Must Account for Quality and Task — prateekj · 2026-08-13