Qwen 3.8-Max Debuts on Modal with 1M Context
Alibaba's Qwen 3.8-Max (Qwen3.8-2.4T-A95B) is now on Modal with a 1M context window and a custom DFlash speculative decoder trained on tool-call data, featuring 2.4T parameters.
2026-08-13 ~ 2026-08-15 · 4 related posts
- Episode 1: Alibaba's Qwen Open-Sources Qwen3.8-Max, a 2.4T-Parameter Flagship MoE Model(2026-08-12, 21 posts)
- Episode 2: Unsloth Shrinks Qwen3.8 by 91% for Local Deployment(2026-08-13, 3 posts)
- Episode 3: vLLM Announces Day-0 Support for Qwen3.8 2.4T Model(2026-08-13, 2 posts)
- Episode 4: Together AI Launches Serverless Inference for Qwen3.8-2.4T-A95B(2026-08-13, 2 posts)
- Episode 5: Qwen 3.8-Max Debuts on Modal with 1M Context(2026-08-13, 4 posts)
- Qwen3.8-2.4T-A95B Available on Modal with 1M Token Context — AAAzzam · 2026-08-13
- Qwen3.8-Max launches on Fireworks with Day-0 support: 2.4T-param MoE for agents and coding — Alibaba_Qwen · 2026-08-14
- Qwen 3.8-Max Now on Modal: 2.4T Params, 1M Context, Custom DFlash Speculator — AAAzzam · 2026-08-14
- Qwen3.8-2.4T-A95B now on SiliconFlow: $2/M input, $6/M output — Alibaba_Qwen · 2026-08-15