DeepSeek releases V4.1 Flash: 552B MoE, 1M context, SGLang day-0 support
BanghuaZ · x · 2026-09-10
DeepSeek officially launched V4.1-Flash, the smallest model in its new architecture family, with native visual understanding, faster inference and higher throughput — weights are open. SGLang and Miles shipped day-0 inference and RL support.
Key specs:
- Compressed KV shared across layers
- Two-stage sparse indexer
- 196B Engram lookup memory
- 552B backbone, 16B active decode / 8B prefill, up to 1M context, natively multimodal
The team promises "very exciting performance upgrades" in the coming days.
Related event: DeepSeek Unveils Open-Source V4.1-Flash MoE Model(28 posts)→
More from Models
- DeepSeek V4.1 Flash is actually 748B params, safetensors analysis shows — DistanceSolar1449 · 2026-09-10
- Kimi K3 lands on RunPod: 2.8T params, 1M context, $3/$15 per 1M tokens — Kimi_Moonshot · 2026-09-10
- Dev burns 300M tokens on GLM 5.3 in a week and still has quota left — saibharadwaj · 2026-09-10
- Follow-up: a 3T-parameter model may already exist, scaling issues remain the wildcard — teortaxesTex · 2026-09-10
- Speculation: DeepSeek V4.1 Pro could be a 3.1T-param MoE with 2.6TB disk footprint — teortaxesTex · 2026-09-10
- New Book Teaches Beginners to Build and Fine-Tune Their Own GPT-Style SLMs, With Colab Notebooks — Roger_M_Taylor · 2026-09-10