DeepSeek V4.1 Flash has just 20 decoder layers, shallower than its predecessor
teortaxesTex · x · 2026-09-10
Dissecting the freshly released DeepSeek V4.1-Flash weights, teortaxesTex found a surprisingly shallow architecture: strictly 20 decoder layers (about 40 total), three layers shallower than V4-Flash, and notably no looped/recursive transformer design despite the recent hype around that idea.
He also estimates the model at roughly 5.6 Sol. The extreme shallowness relative to its performance has sparked community debate over DeepSeek's architectural trade-offs.
More from Models
- Leaked brief: DeepSeek V4.1 Flash at 552B params, foldable iPhone at $1,999, ChatGPT voice limits raised — testingcatalog · 2026-09-10
- CoT similarity test suggests Qwen3.8 may have been trained on GPT 5.5 reasoning traces — Chromix_ · 2026-09-10
- DeepSeek accused of 'pretending linear attention doesn't exist' in new architecture — teortaxesTex · 2026-09-10
- Hands-on with GLM 5.3 Flash: great at coding and research, poor at trading and ideation — ManagementNo5153 · 2026-09-10
- Tell an agent it has 1M token budget and it reasons 3-5x longer: a metacognition experiment — paraschopra · 2026-09-10
- DeepSeek's new open model beats GLM 5.3 and Kimi K3 at 4-10x lower price — deedydas · 2026-09-10