DeepSeek V4.1-Flash leak: 552B params claiming GPT-5.6-level scores via CED architecture
iScienceLuvr · x · 2026-09-10
A third-party post claims DeepSeek released DeepSeek-V4.1-Flash, with benchmarks allegedly matching "GPT-5.6 Sol" despite only 552B params — the poster himself suspects benchmark gaming. Key architecture details cited:
- A Causal Encoder-Decoder (CED) design where decoder global KV is projected from final encoder hidden states, bypassing full decoder KV computation
- Only 8B params active per token during prefill and 16B during decode, aimed at input-heavy agentic workloads
- Described as an SWA-based local-processing backbone augmented with compressed global context, with aggressive KV cache compression as the core focus
Unverified: DeepSeek has not officially confirmed this release.
Related event: Leaked DeepSeek V4.1 Benchmarks Point to New 552B Architecture(15 posts)→
More from Models
- DeepSeek V4.1 Flash is actually 748B params, safetensors analysis shows — DistanceSolar1449 · 2026-09-10
- Kimi K3 lands on RunPod: 2.8T params, 1M context, $3/$15 per 1M tokens — Kimi_Moonshot · 2026-09-10
- Dev burns 300M tokens on GLM 5.3 in a week and still has quota left — saibharadwaj · 2026-09-10
- Follow-up: a 3T-parameter model may already exist, scaling issues remain the wildcard — teortaxesTex · 2026-09-10
- Speculation: DeepSeek V4.1 Pro could be a 3.1T-param MoE with 2.6TB disk footprint — teortaxesTex · 2026-09-10
- New Book Teaches Beginners to Build and Fine-Tune Their Own GPT-Style SLMs, With Colab Notebooks — Roger_M_Taylor · 2026-09-10