FULL STORY
Alibaba's Qwen3.8-Max Release and Ecosystem Adoption
Alibaba released the 2.4T parameter Qwen3.8-Max, quickly followed by Day-0 support and serverless hosting from platforms like vLLM and Together AI.
2026-08-12 ~ 2026-08-13 · 3 episodes · 21 posts
Episode 1 · Alibaba Open-Sources 2.4T Parameter Flagship Qwen3.8-Max (2026-08-12, 17 posts)
Alibaba's Qwen team has debuted its new Qwen3.8-2.4T-A95B model on Hugging Face, swiftly climbing the trending list. The model boasts a total parameter count of 2.4 trillion (2.4T) and utilizes a Mixture of Experts (MoE) architecture, activating approximately 95 billion (95B) parameters per inference.
Confirmed
- The model's page is live on Hugging Face, confirming a total parameter count of 2.4T with 95B active parameters.
- It has officially hit the Hugging Face trending list.
- The model will feature open weights, with Nebius Token Factory announced as a Day 0 launch partner.
Unconfirmed
- Based on leaked configuration details viewed by developer @multimodalart, the model natively supports robust vision and Agent tool-calling capabilities. However, these feature specifics currently stem solely from the leaked data.
Why it matters
- As the latest flagship in the Qwen series, its massive 2.4T parameter scale makes it one of the industry's most closely watched ultra-large models. Its open-weights approach and native multimodal/Agent capabilities have generated significant buzz across the AI community.
- Qwen Releases Qwen3.8-2.4T-A95B Model on Hugging Face — de4dee · 2026-08-12
- Qwen3.8-2.4T Model Surfaces on Hugging Face with Agent Capabilities — multimodalart · 2026-08-12
- Qwen's New 2.4T MoE Model Tops Hugging Face Trending — Qwen · 2026-08-12
- Alibaba's Qwen3.8 Goes Open Weight: 2.4T Parameters with 95B Active — Arindam_1729 · 2026-08-12
- Alibaba's Qwen3.8-2.4T-A95B Drops, Autonomously Optimizes Inference Engine — BanghuaZ · 2026-08-12
- Qwen3.8-Max Weights Released: 2.4T Total Parameters — Yuchenj_UW · 2026-08-12
- Qwen3.8-Max Weights Released with Commercial Revenue Caps — cedric_chee · 2026-08-13
- Alibaba releases Qwen3.8 Max: 2.4T-param MoE, 1M context, open for commercial use — AdinaYakup · 2026-08-13
- Alibaba releases Qwen3.8 Max: 2.4T-param MoE, 1M context, open for commercial use — AdinaYakup · 2026-08-13
- Qwen Releases Qwen3.8-Max: 2.4T Parameters with 1M Context natively — Scobleizer · 2026-08-13
- Qwen3.8 2.4T MoE FP8 Quantization Hits Hugging Face Trending — Qwen · 2026-08-13
- Alibaba's Qwen Open-Sources Qwen3.8-2.4T-A95B Model — xiaosun86 · 2026-08-13
- Alibaba Releases 2.4T-Parameter Open-Weight Model Qwen3.8-Max — NVIDIAAI · 2026-08-13
- Alibaba Launches 2.4T Parameter Qwen3.8-Max with 1M Context and Multimodal Support — togethercompute · 2026-08-13
- Qwen3.8-2.4T-A95B now live on Together AI: 2.4T params, 256K context for coding and agents — togethercompute · 2026-08-13
- Alibaba Releases Qwen3.8-Max: 2.4T Parameter MoE with Native Multimodal Support — baseten · 2026-08-13
- Alibaba Qwen Open-Sources Flagship Qwen3.8-2.4T-A95B with Trillion Parameters — 智东西 · 2026-08-13
Episode 2 · vLLM Announces Day-0 Support for Qwen3.8 2.4T Model (2026-08-13, 2 posts)
vLLM has announced Day-0 support for Alibaba's massive Qwen3.8-2.4T-A95B open-source model, providing out-of-the-box 4-bit quantization to enable single-node execution.
- vLLM Announces Day-0 Support for Qwen3.8 2.4T with Ready-to-use 4-bit Checkpoints — vllm_project · 2026-08-13
- vLLM Day-0 Support for Qwen3.8-2.4T-A95B: Runs on Single NVIDIA/AMD Nodes — vllm_project · 2026-08-13
Episode 3 · Together AI Launches Serverless Inference for Qwen3.8-2.4T-A95B (2026-08-13, 2 posts)
Together AI has launched serverless inference support for the Qwen3.8-2.4T-A95B model, featuring 256K context and optimized for high-throughput coding and agentic workloads, enabling robust long-range task management.
- Qwen3.8-2.4T-A95B for long-running agents: 256K context, self-testing, function calling — togethercompute · 2026-08-13
- Together AI launches serverless inference for Qwen3.8-2.4T-A95B with 99.9% SLA — togethercompute · 2026-08-13