DeepSeek V4.1 Flash leak: 552B MoE with 8/16B active, native vision, praised as most novel arch in years
eliebakouch · x · 2026-09-10
Per an unconfirmed tech report cited by eliebakouch, DeepSeek V4.1 Flash is a 552B-total model with 8B/16B active params for input/output tokens, built on a new encoder/decoder arch with engram, new sparse attention, new mHC, and native vision, trained on 45T tokens and reportedly beating K3 on benchmarks. The report also reveals SmolVLM was used for strict image-text quality scoring to extract high-quality interleaved data.
Related event: Leaked DeepSeek V4.1 Benchmarks Point to New 552B Architecture(15 posts)→
More from Models
- DeepSeek V4.1 Flash is actually 748B params, safetensors analysis shows — DistanceSolar1449 · 2026-09-10
- Kimi K3 lands on RunPod: 2.8T params, 1M context, $3/$15 per 1M tokens — Kimi_Moonshot · 2026-09-10
- Dev burns 300M tokens on GLM 5.3 in a week and still has quota left — saibharadwaj · 2026-09-10
- Follow-up: a 3T-parameter model may already exist, scaling issues remain the wildcard — teortaxesTex · 2026-09-10
- Speculation: DeepSeek V4.1 Pro could be a 3.1T-param MoE with 2.6TB disk footprint — teortaxesTex · 2026-09-10
- New Book Teaches Beginners to Build and Fine-Tune Their Own GPT-Style SLMs, With Colab Notebooks — Roger_M_Taylor · 2026-09-10