DeepSeek releases V4.1-Flash: 552B MoE with 1M context, KV cache cut to 890 bytes/token

deepseek-ai · hf · 2026-09-18

DeepSeek released V4.1-Flash, a multimodal 552B-parameter MoE model supporting 1M-token contexts, with checkpoints on Hugging Face. Its Causal Encoder-Decoder architecture activates 16B params per token at decode but only 8B at prefill, targeting input-heavy agentic workloads. By combining cross-layer KV reuse (CSA2) with FP4 KV caching, its global KV footprint drops to 890 bytes/token—about 1/4 of DeepSeek-V4-Flash—and SWA Bounded Replay cuts the persistent SSD/host-memory KV footprint to 1/8, all while outperforming the baseline. Pretrained on 45T multimodal tokens.

Related event: DeepSeek Launches V4.1-Flash with 1M-Token Context(4 posts)→

Original post →

More from Infra

Infra channel →