Qwen 3.8-Max Now on Modal: 2.4T Params, 1M Context, Custom DFlash Speculator
AAAzzam · x · 2026-08-14
Alibaba's Qwen 3.8-Max is now available on Modal, with full 1M context window and a custom DFlash speculator trained on tool-call-heavy data. The model has 2.4T parameters (95B active) and is a Modal call away.
Related event: Qwen 3.8-Max Debuts on Modal with 1M Context(2 posts)→
More from Infra
- Qwen3.8-27B-FP8 on GH200: 10 concurrent streaming requests, first token in 10ms — MaziyarPanahi · 2026-08-15
- Qwen3.8-27B Serving Configs: DGX Spark vLLM and RTX 4090 llama.cpp — erdaltoprak · 2026-08-15
- AI video billing metered by pixel area, not duration: developer warns to reconcile invoices — DryProgress9179 · 2026-08-15
- Analyst: Memory Cycle Structural, Sentiment Should Shift — BenBajarin · 2026-08-15
- Attention-FFN disaggregation: a new direction for inference speedups? — tokenbender · 2026-08-14
- DeepSeek-V4-Pro launches: 1.6T-param MoE cuts inference FLOPs to 27% of V3.2 — AccBalanced · 2026-08-14