Qwen3.8 Open-Sources 2.4T Parameter MoE Flagship Model
togethercompute · x · 2026-08-14
Alibaba's Qwen team has launched their new open-weight flagship model, Qwen3.8-2.4T-A95B, now available on Together AI.
The model uses a sparse Mixture-of-Experts (MoE) architecture with 2.4 trillion total parameters and 95 billion active parameters. Key features include:
- Multimodal Support: Qwen's first multimodal model above one trillion parameters, processing both text and images.
- Long Context: Supports a 256K context window with up to 128K output tokens.
- Coding & Agentic Focus: Optimized for coding and long-horizon agentic tasks.
- Adjustable Reasoning: Features always-on thinking with adjustable effort levels across low, high, and xhigh.
Related event: Alibaba Open-Sources 2.4T Parameter Flagship Model Qwen3.8-Max(21 posts)→
More from Models
- Anthropic's Suspicious Silence Hints at Upcoming 'Claudette' Drop — beffjezos · 2026-08-14
- DeepSeek V4-Pro slows to crawl on long projects; reloading reveals it's further ahead — teortaxesTex · 2026-08-14
- Google Focuses on Smaller, Faster AI Models Leveraging Massive Search Scale — haider1 · 2026-08-14
- Notion Launches Knowledge Board: Evaluating LLMs on Real-World Traffic Instead of Benchmarks — ivanhzhao · 2026-08-14
- DeepSeek V4 Pro Hits Baseten APIs: 1.7T Parameters, MIT License — baseten · 2026-08-14
- Grok Blocks Intimate Likeness Edits, Proving Musk's 'Legal = Allowed' Wrong — firasd · 2026-08-14