Intern-S2-397B: Shanghai AI Lab's 397B-A17B scientific MoE lands day-0 in vLLM

vllm_project · x · 2026-09-14

Shanghai AI Laboratory released Intern-S2-397B, a multimodal MoE foundation model for long-horizon scientific research: 397B total / 17B active params across 512 experts on the Qwen3.5 hybrid linear/full-attention architecture, 262K context, and a built-in shared-weight MTP head for speculative decoding.

Key claims:

Deployment: day-0 vLLM support (vLLM ≥ 0.22.1). The official FP8 checkpoint runs on 8×H100/H200 with TP=8; BF16 needs 8×H200. Requires trust-remote-code (custom InternS1Tokenizer), DeepGEMM disabled, FlashInfer TRTLLM MoE backend on NVIDIA. Time-series inference is LMDeploy-only for now.

Related event: Shanghai AI Lab Open-Sources Intern-S2-397B with Multimodal and Agent Powers(3 posts)→

Original post →

More from Infra

Infra channel →