DeepSeek previews new architecture: 552B MoE with just 8B active input parameters

zephyr_z9 · x · 2026-09-10

DeepSeek officially previewed its next-gen model: an asymmetric 'more intelligence, less cost' Causal Encoder–Decoder MoE with 552B total parameters but only 8B active for input and 16B for output. New pre-training methods plus larger-scale RL post-training reportedly deliver benchmark results ahead of flagship models including DeepSeek-V4-Pro. Details to follow in a 6-part thread.

Related event: DeepSeek Unveils Open-Source V4.1-Flash MoE Model(28 posts)→

Original post →

More from Models

Models channel →