Qwen3.8-2.4T-A95B open weights land on AWS: single 8×B300 node with vLLM

AWS ML Blog · rss · 2026-09-10

Alibaba's Qwen team released Qwen3.8-2.4T-A95B on August 12, 2026 — the first open-weights Qwen-Max-class model. An AWS ML Blog post walks through deploying it on SageMaker HyperPod with vLLM.

Architecture highlights

Deployment

Vendor benchmarks show strengths in PaperBench (93.0), IFBench (82.8), and terminal coding (86.6), comparable to frontier models, with headroom on repository-level tasks (SWE-bench Pro) — a credible self-hosted option for coding agents and research pipelines.

Original post →

More from Infra

Infra channel →