SGLang Hits 873 tok/s on 4×GB300 With DeepSeek V4.1 Flash Day-0 Support
hsu_byron · x · 2026-09-12
SGLang shipped Day-0 support for DeepSeek V4.1 Flash and pushed performance to 873 tok/s at BS=1 on 4×GB300 within 24 hours, via FP8 GEMM fast paths, kernel fusion and overlap, DSpark optimization, and MoE TP4.
V4.1 Flash introduces major architectural changes — a causal encoder-decoder design, CSA2 shared KV, Engram, mHC, and DSpark — all supported from day zero. The team published a hands-on engineering guide with step-by-step reproduction instructions; DeepGEMM and FlashMLA optimizations are coming next.
Related event: SGLang delivers Day 0 support for DeepSeek V4.1 Flash, hitting 873 tok/s(2 posts)→
More from Infra
- Full Talk Slides Released: How Inference Engines Actually Work, End to End — zainhas · 2026-09-12
- Full "How Inference Engines Actually Work" Slides Now Live on Google Slides — zainhas · 2026-09-12
- Zilliz CTO: Agent memory is a long-lived systems problem, not an index feature — J_Luan_ · 2026-09-12
- Draw Things update adds MiniMax H3 with LoRA/TeaCache and Krea 2 model imports — antirez · 2026-09-12
- Wafer launches 'most comprehensive' AI performance engineering repo, starting with Transformer inference deep-dive — ycombinator · 2026-09-12
- Together serves 23%-30% of all OpenRouter traffic for GLM 5.3 models — zhyncs42 · 2026-09-12