Intern-S2-397B: Shanghai AI Lab's 397B-A17B scientific MoE lands day-0 in vLLM
vllm_project · x · 2026-09-14
Shanghai AI Laboratory released Intern-S2-397B, a multimodal MoE foundation model for long-horizon scientific research: 397B total / 17B active params across 512 experts on the Qwen3.5 hybrid linear/full-attention architecture, 262K context, and a built-in shared-weight MTP head for speculative decoding.
Key claims:
- Top-tier open-source general ability in knowledge, coding, and agents
- Strong on scientific benchmarks like Biology-Instructions and Mol-Instructions
- Leading open-source results on IMO-Proof and AdvancedMathBench, claimed at the level of top closed models such as Gemini 3.1 Pro
Deployment: day-0 vLLM support (vLLM ≥ 0.22.1). The official FP8 checkpoint runs on 8×H100/H200 with TP=8; BF16 needs 8×H200. Requires trust-remote-code (custom InternS1Tokenizer), DeepGEMM disabled, FlashInfer TRTLLM MoE backend on NVIDIA. Time-series inference is LMDeploy-only for now.
More from Infra
- llama.cpp Adds Maple 20B-A1B Ternary MoE Architecture for CPU and Low-VRAM Devices — jacek2023 · 2026-09-14
- Claude usage boosts quietly removed, fueling talk that 'The Great Compute Crunch has begun' — jacob_posel · 2026-09-14
- Unions urged to halt AI datacenter buildout until jobs and grid use are protected — nordicinst · 2026-09-14
- SK hynix completes HBM4 internal qualification, ushering in custom base die competition — blaizedsouza · 2026-09-14
- One architectural change cuts KV cache 8x: how GQA works, explained with Llama 3 70B — blaizedsouza · 2026-09-14
- A complete breakdown of HBM system architecture, from DDR roots to GDDR7, PIM and HBF alternatives — blaizedsouza · 2026-09-14