DeepSeek launches V4.1-Flash: 552B MoE with 8B active params, 1/4 the KV cache
thione · x · 2026-09-21
DeepSeek released DeepSeek-V4.1-Flash, the smallest model in its new architecture family: a 552B-parameter MoE with an asymmetric Causal Encoder–Decoder design, activating just 8B parameters for input and 16B for output, with native visual understanding.
Key points:
- KV cache needs cut to 1/4 the HBM and 1/8 the SSD storage versus the previous generation, significantly lowering agent cache-hit costs;
- New pretraining methods plus larger-scale RL post-training reportedly beat flagship V4-Pro; DeepSeek is phasing out V4-Pro (all v4-pro requests route to V4.1-Flash rates from Sept 14);
- Live on the API as deepseek-flash with peak/off-peak pricing (off-peak at 50%);
- Partners WorkBuddy (incl. CodeBuddy) and OpenCode fully support it.
Related event: DeepSeek Unveils V4.1-Flash, a 552B MoE Redesigned for Agents(3 posts)→
More from Infra
- Meta partners with Arm on Arm AGI CPU, its first AI-era data center CPU — bookwormengr · 2026-09-21
- Starlink is becoming core infrastructure for rural education across Latin America — XFreeze · 2026-09-21
- HF speech-to-speech merges NVIDIA Parakeet Unified STT backend — andimarafioti · 2026-09-21
- Apple wins the AI infra lottery: M5 Ultra packs 512GB unified memory for local AI — Hesamation · 2026-09-21
- Z Image Turbo runs locally on a Snapdragon 8 Elite phone at 40s/image — Fine_Philosopher_882 · 2026-09-21
- Linux on 8GB RAM: you can run LLMs and do real development — gnukeith · 2026-09-21