FULL STORY

DeepSeek V4.1-Flash: Launch and Real-World Tests

DeepSeek launched V4.1-Flash, a 552B MoE model with million-token context aimed at long-horizon agents. Developers then reported cutting AI costs by 90% after large-scale testing.

2026-09-18 ~ 2026-09-19 · 2 episodes · 6 posts

Episode 1 · DeepSeek Flash Becomes Default Workhorse, Cutting AI Costs 90% (2026-09-18, 2 posts)

Developer @btsouth pushed billions of tokens through DeepSeek V4.1 Flash in agent workflows and now treats it as a default workhorse for coding and debugging, cutting AI spend from about $300 to $30 per month.

Episode 2 · DeepSeek Launches V4.1-Flash with 1M-Token Context (2026-09-18, 4 posts)

DeepSeek released V4.1-Flash, a 552B-parameter multimodal MoE (8B active) with native vision and 1M-token context, compressing KV cache to 890 bytes per token; it is now live on Inco (topping speed charts at 532 tokens/s) and Nebius Token Factory.