DeepSeek V4.1 Flash weights out: 552B multimodal, 1M context, SGLang day-0 support
zhaoran_wang · x · 2026-09-11
SGLang announced day-0 inference and RL support for DeepSeek V4.1 Flash, whose weights are now out. V4.1 extends the V4 stack with compressed KV shared across layers, a two-stage sparse indexer, and a 196B Engram lookup memory.
It's natively multimodal with a 552B backbone (16B active decode, 8B prefill) and supports up to 1M context, with 'very exciting performance upgrades' teased for the coming days. The retweet notes Chinese open-source release cycles have shrunk from yearly to monthly.
Related event: DeepSeek releases V4.1-Flash: 552B MoE beats flagships at low cost(56 posts)→
More from Models
- ApprenticeBench: closed model scores 72% vs open Kimi K3 at 18% on real jobs — ysu_nlp · 2026-09-11
- CursorBench 4.0 launches; Muse Spark 1.3 matches Sol at under 40% the cost — jyangballin · 2026-09-11
- Sparse attention as multilevel retrieval: the DeepSeek V3.2 trick mirrors search ranking stacks — nptacek · 2026-09-11
- muse spark 1.3 scores strong and cheap on CursorBench 4 — infoxiao · 2026-09-11
- Users hope Haiku 3.5 returns as model cutoffs draw criticism — repligate · 2026-09-11
- kalomaze: 'Fable 5' Partly Suffers from Undercooked Post-Training — kalomaze · 2026-09-11