DeepSeek unveils asymmetric Causal Encoder-Decoder: 552B MoE with just 8B active input params
ChengleiSi · x · 2026-09-10
DeepSeek officially announced a new asymmetric architecture: a 552B-parameter MoE with a novel Causal Encoder–Decoder design using only 8B active parameters for input and 16B for output. New pre-training methods plus larger-scale RL post-training deliver benchmark results ahead of flagship models including DeepSeek-V4-Pro. Researchers note steady ProgramBench gains.
More from Models
- NeoHorse-1-4B, a Qwen3.5-based agentic model, trends on Hugging Face — TokenRhythm · 2026-09-10
- Chinese model 3D face-off: DeepSeek V4.1 Flash crushes Kimi K3 and GLM-5.3 — teortaxesTex · 2026-09-10
- DeepSeek 4.1 Flash spotted running the Boeing bench, results pending — victormustar · 2026-09-10
- Sentry CEO: switch off the priciest reasoning-tier models — you won't notice a performance difference — zeeg · 2026-09-10
- Latest Models Are Now Surprisingly Good at Driving ffmpeg — Flomerboy · 2026-09-10
- Astra and Fable 5.1 benchmarks barely overlap, making leaderboard comparisons misleading — recro69 · 2026-09-10