DeepSeek releases V4.1-Flash: 552B MoE with 1M context, KV cache cut to 890 bytes/token
deepseek-ai · hf · 2026-09-18
DeepSeek released V4.1-Flash, a multimodal 552B-parameter MoE model supporting 1M-token contexts, with checkpoints on Hugging Face. Its Causal Encoder-Decoder architecture activates 16B params per token at decode but only 8B at prefill, targeting input-heavy agentic workloads. By combining cross-layer KV reuse (CSA2) with FP4 KV caching, its global KV footprint drops to 890 bytes/token—about 1/4 of DeepSeek-V4-Flash—and SWA Bounded Replay cuts the persistent SSD/host-memory KV footprint to 1/8, all while outperforming the baseline. Pretrained on 45T multimodal tokens.
Related event: DeepSeek Launches V4.1-Flash with 1M-Token Context(4 posts)→
More from Infra
- First Cafe Compute Nairobi meetup demos Cerebras API, calls for an /explain interpretability endpoint — paw_lean · 2026-09-19
- Cerebras launches Money Agent, a Qwen 3 27B-powered personal finance assistant — Alibaba_Qwen · 2026-09-19
- GPU host warns: renter exploited his rig for attacks, Clore.AI blocked him for reporting it — anomaly256 · 2026-09-19
- 3M paid $10.3B to quit PFAS — AI data centers just made it a growth market again — aakashgupta · 2026-09-19
- Random KV Cache eviction rivals top baselines and boosts vLLM throughput 32-43% — 机器之心 · 2026-09-19
- $250 of modded mining cards, 30GB VRAM: old i7 PC runs Qwen at 30 tok/s with patched drivers — HFq_Dev · 2026-09-19