DeepSeek releases V4.1-Flash: 552B MoE claimed to beat V4-Pro on cost and speed
eyishazyer · x · 2026-09-11
DeepSeek officially released V4.1-Flash on Sept. 10: a 552B-parameter MoE with a new causal encoder-decoder split (8B active params for input, 16B for output). DeepSeek claims its KV cache uses a quarter of the HBM and an eighth of the SSD footprint of the previous generation, and that Flash beats V4-Pro on performance, cost, speed and runtime — enough that all V4-Pro API traffic will be routed to Flash from Sept. 14. Off-peak cache-hit input pricing is reportedly dropping to RMB 0.02 per million tokens, per TechNode, though unconfirmed on DeepSeek's pricing page.
More from Infra
- Training a 6-Expert MoE GPT-2 From Scratch on a Single RTX 3090 in 8 Days — rasbt · 2026-09-11
- B3IQ Sells Eight Figures of GPUs in Two Weeks, Bets AI Infra Is a $100B Market — templecrash · 2026-09-11
- It Cost $100 in API Credits for an AI Agent to Install Free Software — MartinGTobias · 2026-09-11
- bartowski unveils per-tensor layout maps for GGUF quantization, tests show across-the-board gains — noneabove1182 · 2026-09-11
- antirez runs DeepSeek v4.1 Flash locally on a 128GB M5 Max, SSD streaming surprisingly fast — antirez · 2026-09-11
- OpenAI could 7x its training compute tomorrow: why open-source models still trail by one generation — soumitrashukla9 · 2026-09-11