SGLang hits 873 tok/s on DeepSeek V4.1 Flash within 24 hours of launch
BanghuaZ · x · 2026-09-12
SGLang shipped Day 0 support for DeepSeek V4.1 Flash and pushed performance to 873 tok/s at BS=1 on 4×GB300 within 24 hours.
- Optimizations include FP8 GEMM fast paths, kernel fusion and overlap, DSpark optimization, and MoE TP4
- V4.1 Flash introduces major architecture changes: causal encoder-decoder design, CSA2 shared KV, Engram, mHC, and DSpark — all supported from day one
- A hands-on engineering guide with step-by-step reproduction instructions has been published
- DeepGEMM, FlashMLA, and more optimizations are coming next
Related event: SGLang delivers Day 0 support for DeepSeek V4.1 Flash, hitting 873 tok/s(2 posts)→
More from Infra
- The Global Race for Cheap Power: Where AI Data Centers Should Actually Go — pravchaw · 2026-09-12
- VCs float 'hardware revenue derivative': fund compute costs via revenue share, not equity — ns123abc · 2026-09-12
- Relace hits 1T tokens/day on OpenRouter, serving 37% of DeepSeek v4 Flash traffic — stuffyokodraws · 2026-09-12
- Yutori's Navigator n2 runs browser agents at $1.46 per task on OSWorld 2.0 vs $13-$40+ for frontier models — DhruvBatra_ · 2026-09-12
- Auto-derived FlashAttention with SMEM and tensor core assignment shown off — vtabbott_ · 2026-09-12
- AI progress timing debate: same-node hardware gains deliver a one-time compute windfall — cis_female · 2026-09-12