DeepSeek Unveils V4.1-Flash, a 552B MoE Redesigned for Agents
DeepSeek released V4.1-Flash, its smallest new-architecture model: a 552B-parameter MoE with only 8B active, featuring an asymmetric Causal Encoder-Decoder design and KV cache compressed to 890 bytes per token for agent workloads.
2026-09-21 ~ 2026-09-21 · 3 related posts
- DeepSeek-V4.1-Flash redesigns the Transformer for agents, cutting KV cache to 890 bytes/token — AndLukyane · 2026-09-21
- Recap claims Google shipped Gemini 3.8 Live and DeepSeek released V4.1-Flash MoE — thione · 2026-09-21
- DeepSeek launches V4.1-Flash: 552B MoE with 8B active params, 1/4 the KV cache — thione · 2026-09-21