Rumored DeepSeek V4.1 Flash details point to asymmetric-activation MoE with much lower cost

gaganghotra_ · x · 2026-09-10

A viral repost claims unverified details about DeepSeek V4.1 Flash: a 552B MoE with a novel Causal-Encoder-Decoder architecture, asymmetric activation (8B reading, 16B writing) for much lower cost, plus new pretraining and large-scale RL posttraining allegedly beating V4 Pro and other flagships. Note these figures conflict with official/vLLM specs (769B total / 15.5B active) — treat with caution.

Related event: Alleged DeepSeek V4.1 and V4.1 Flash Benchmarks and Architecture Leak(12 posts)→

Original post →

More from Models

Models channel →