Alibaba Qwen Open-Sources Flagship Qwen3.8-2.4T-A95B with Trillion Parameters
智东西 · wechat · 2026-08-13
Alibaba's Qwen team open-sourced the weights for its flagship model, Qwen3.8-2.4T-A95B, featuring 2.4 trillion total parameters with 95 billion activated per token. It natively supports over 1 million tokens of context.
Capabilities & Benchmarks
The model excels in coding and long-horizon agentic tasks, topping the PaperBench and OSWorld benchmarks. However, it trails slightly behind GPT-5.6Sol and Fable5 on TerminalBench2.1 and SWE-benchPro. Qwen demonstrated the model's autonomous capabilities by having it code continuously for 16 days and manage a virtual e-commerce business.
Deployment & Pricing
The model defaults to a mandatory thinking mode. UnslothAI managed to compress the model from 4.9TB to 397GB using dynamic quantization, significantly lowering local deployment barriers. Overseas API pricing is set at $2 per million input tokens and $6 per million output tokens.
Related event: Alibaba Open-Sources 2.4T Parameter Flagship Qwen3.8-Max(17 posts)→
More from Models
- DeepSeek Cybersecurity Test: Top Recall but Bottom Precision — teortaxesTex · 2026-08-13
- Building 'The Office' Agent Simulation with Grok 4.6: A Major Leap in Speed and Capability — mattyp · 2026-08-13
- Sakana AI Updates Chat with New Fugu Model and Code Execution — SakanaAILabs · 2026-08-13
- Gemini V4-Pro Disappoints in Tests, Suspected to Be Hampered by Internal Distillation — teortaxesTex · 2026-08-13
- DeepSeek V4-Pro Ranks #2 Open-Weight Model, Accused of Relying on pass@2 — teortaxesTex · 2026-08-13
- GPT Models Struggle with Pixel Art: UI Interaction Fails and Poor Generation — breath_mirror · 2026-08-13