DeepSeek V4.1 Flash has just 20 decoder layers, shallower than its predecessor

teortaxesTex · x · 2026-09-10

Dissecting the freshly released DeepSeek V4.1-Flash weights, teortaxesTex found a surprisingly shallow architecture: strictly 20 decoder layers (about 40 total), three layers shallower than V4-Flash, and notably no looped/recursive transformer design despite the recent hype around that idea.

He also estimates the model at roughly 5.6 Sol. The extreme shallowness relative to its performance has sparked community debate over DeepSeek's architectural trade-offs.

Original post →

More from Models

Models channel →