DeepSeek V4.1 Flash at 552B params beats its 1.6T flagship while costing 4x less

ArtificialAnlys · x · 2026-09-11

DeepSeek released V4.1 Flash, a 552B-parameter model with a new causal encoder-decoder architecture (8B active input / 16B active output params) that overtakes the 1.6T V4 Pro 0813 as the company's flagship, scoring 40 on the AA Intelligence Index.

Pricing: $0.30/$1.20 per 1M input/output tokens, cached input at $0.006/1M (98% discount), off-peak pricing another 50% off — 20% cheaper than the previous Flash and 4x cheaper than V4 Pro.

Key results:

Caveat: extremely verbose — 89k output tokens per Index task, 62% more than its own Pro model and more than frontier models like Fable 5.1. Still only $0.27 per task, 7x cheaper than GLM-5.3/Kimi K3.

Related event: DeepSeek Releases Open-Source V4.1 Flash: 552B MoE Undercuts and Outperforms Flagships(55 posts)→

Original post →

More from Models

Models channel →