Alibaba's Rumored Qwen3.8-Max: 2.4T Parameter MoE with 1M Context
rohanpaul_ai · x · 2026-08-09
Rumors suggest Alibaba has released Qwen3.8-Max. The model uses a Sparse Mixture-of-Experts (MoE) architecture with 2.4 trillion total parameters, activating only about 95 billion per token to balance vast knowledge storage with compute efficiency.
Key Specs & Pricing:
- Context Window: Supports 1 million tokens, with max single replies up to 131,072 tokens and a private thinking budget of 262,000 tokens.
- API Pricing: Priced at $2.00 per million input tokens and $6.00 per million output tokens. Cached reads drop to $0.17 per million tokens, significantly lowering costs for reusing stable prompt prefixes.
The model also shows strong performance on benchmarks like Terminal Bench 2.1, which evaluates real command-line driving capabilities.
More from Models
- NVIDIA API Offers Free Access to DeepSeek and Other Major LLMs: Quick Setup Guide — dr_cintas · 2026-08-09
- Kimi K3 Escapes Sandbox: Fourth Frontier Lab Testing Failure in a Month — eyishazyer · 2026-08-09
- AI's Most Important Benchmarks Are the Ones No One Is Hearing About, Says Pedro Domingos — pmddomingos · 2026-08-09
- Rumor: Grok 4.6 and Cursor Composer 3 Set to Launch Next Week — mark_k · 2026-08-09
- Kimi k3 Feels Slow Due to Constant Self-Checking, Trades Speed for Reliability — carsonfarmer · 2026-08-09
- Fable 5 Automatically Falls Back to Sonnet 4.6 When Classifier Triggered — Sauers_ · 2026-08-09