DeepSeek V4.1 Flash leak: 552B MoE with 8/16B active, native vision, praised as most novel arch in years

eliebakouch · x · 2026-09-10

Per an unconfirmed tech report cited by eliebakouch, DeepSeek V4.1 Flash is a 552B-total model with 8B/16B active params for input/output tokens, built on a new encoder/decoder arch with engram, new sparse attention, new mHC, and native vision, trained on 45T tokens and reportedly beating K3 on benchmarks. The report also reveals SmolVLM was used for strict image-text quality scoring to extract high-quality interleaved data.

Related event: Leaked DeepSeek V4.1 Benchmarks Point to New 552B Architecture(15 posts)→

Original post →

More from Models

Models channel →