DeepSeek V4.1-Flash turns heads: GPT-5.6-level benchmarks at 552B params

teortaxesTex · x · 2026-09-10

DeepSeek released V4.1-Flash, with benchmarks reportedly at GPT-5.6 Sol level despite only 552B parameters, prompting "benchmarkmaxxed" skepticism. The architecture focuses on aggressive KV cache compression: a Causal Encoder-Decoder (CED) design projects decoder global KV from final encoder hidden states, activating just 8B params per token during prefill and 16B during decode—cost-effective for input-heavy agentic workloads. giffmana notes the native multimodal stack still uses SigLIP, suggesting a conservative choice.

Related event: Leaked DeepSeek V4.1 / V4.1 Flash Benchmarks Claim to Rival GPT-5.6 Sol at 552B(19 posts)→

Original post →

More from Models

Models channel →