DeepSeek Ships V4.1-Flash: 552B MoE With 8B Active, 75% Less KV Cache Memory

mark_k · x · 2026-09-11

DeepSeek released DeepSeek-V4.1-Flash, the smallest model in a new architecture family, claiming it beats V4-Pro on performance, cost and speed. It's live on the API and open-sourced on Hugging Face, with V4-Pro requests switching over September 14 ahead of a future V4.1-Pro.

Key specs and efficiency gains:

markk notes: if this is the smallest model in the family, the upcoming V4.1-Pro should be significant.

Related event: DeepSeek open-sources V4.1 Flash: 552B MoE beats flagships at fraction of cost(50 posts)→

Original post →

More from Infra

Infra channel →