Alibaba Releases 2.4T-Parameter Open-Weight Model Qwen3.8-Max
NVIDIAAI · x · 2026-08-13
Alibaba released Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model to date, featuring 2.4T total parameters with 95B activated per token.
- Architecture: A fine-grained Mixture of Experts (MoE) architecture combining full and linear attention, supporting up to a 1M token context window and 128K output length, optimized for reasoning and agentic workloads.
- Inference Performance: According to an NVIDIA technical blog, the model achieves over 4K tokens/second per GPU and 350+ tokens/second per user on the NVIDIA GB300 NVL72 in FP8 precision out of the box.
- Infrastructure: Deploying a 2.4T model requires data-center-scale compute. NVIDIA is collaborating with the open-source ecosystem to provide optimized multi-node deployment recipes.
Related event: Alibaba Releases Qwen3.8-Max: 2.4T-Parameter Open-Weight Model(13 posts)→
More from Infra
- Nvidia Partners with Wall Street to Mobilize $500B for AI Infrastructure — fortune · 2026-08-13
- New Quantization Framework to Run 1.6T Models on a Single B300 GPU at 50 tok/s — dosco · 2026-08-13
- Menlo Park Daytime Electricity Hits 53.8¢/kWh: Running Own GPUs Becomes Irrational — generativist · 2026-08-13
- Running Krea 2 Turbo on 8GB VRAM: RTX 3070 Ti Local Test — niechta · 2026-08-13
- Fluidstack Visits NYSE to Discuss US AI Infrastructure Investment — MxMnr · 2026-08-13
- Open-Source mlx-dspark Boosts LLM Inference on Mac by 3.3x — A-Rahim · 2026-08-13