Alibaba's Rumored Qwen3.8-Max: 2.4T Parameter MoE with 1M Context

rohanpaul_ai · x · 2026-08-09

Rumors suggest Alibaba has released Qwen3.8-Max. The model uses a Sparse Mixture-of-Experts (MoE) architecture with 2.4 trillion total parameters, activating only about 95 billion per token to balance vast knowledge storage with compute efficiency.

Key Specs & Pricing:

The model also shows strong performance on benchmarks like Terminal Bench 2.1, which evaluates real command-line driving capabilities.

Original post →

More from Models

Models channel →