DeepSeek Ships V4.1-Flash: 552B MoE With 8B Active, 75% Less KV Cache Memory
mark_k · x · 2026-09-11
DeepSeek released DeepSeek-V4.1-Flash, the smallest model in a new architecture family, claiming it beats V4-Pro on performance, cost and speed. It's live on the API and open-sourced on Hugging Face, with V4-Pro requests switching over September 14 ahead of a future V4.1-Pro.
Key specs and efficiency gains:
- 552B-parameter MoE with only 8B active for input processing and 16B for output
- Native visual understanding
- KV cache uses just ¼ the GPU memory and ⅛ the SSD storage of the previous generation — a big deal for agents on long tasks
markk notes: if this is the smallest model in the family, the upcoming V4.1-Pro should be significant.
More from Infra
- Positron AI raises $230M Series B at over $1B valuation with Arm backing — seanmcdonaldxyz · 2026-09-11
- Cerebras Fast Inference Flips Agent Workflows: Fewer Parallel Agents, Same Output — MatthewBerman · 2026-09-11
- Baseten acquires Blaxel to build integrated cloud infrastructure for AI agents — baseten · 2026-09-11
- Hyperscalers could factor RSA-1024 for about $30M per number, analysis claims — rickasaurus · 2026-09-11
- Qualcomm's Next Hexagon NPU: 50% More Shared Memory, 30B MoE Models on a Phone — ryanshrout · 2026-09-11
- Vercel Cut CDN P99 Metadata Lookup Latency by 91% Across 80M Route Decisions/sec — cramforce · 2026-09-11