DeepSeek-V4.1-Flash leaks: 552B backbone activating just 8B, 8x less KV cache

bodonoghue85 · x · 2026-09-10

SemiAnalysis congratulated DeepSeek on releasing DeepSeek-V4.1-Flash (third-party, unconfirmed): a 552B causal encoder-decoder activating only 8B params at prefill and 16B at decode, a 196B Engram memory accessed via sparse lookups, and 8x less persistent KV cache than V4-Flash through bounded replay.

Related event: DeepSeek V4.1 Flash Leak: 552B Asymmetric MoE Reportedly Rivals GPT-5.6(20 posts)→

Original post →

More from Models

Models channel →