DeepSeek unveils asymmetric Causal Encoder-Decoder: 552B MoE with just 8B active input params

ChengleiSi · x · 2026-09-10

DeepSeek officially announced a new asymmetric architecture: a 552B-parameter MoE with a novel Causal Encoder–Decoder design using only 8B active parameters for input and 16B for output. New pre-training methods plus larger-scale RL post-training deliver benchmark results ahead of flagship models including DeepSeek-V4-Pro. Researchers note steady ProgramBench gains.

Original post →

More from Models

Models channel →