New model's architecture is 'vanilla': SWA plus MoE with no shared experts, unlike DeepSeek

nrehiew_ · x · 2026-09-23

nrehiew reviews a new model technical report: the architecture is surprisingly simple and vanilla—sliding window attention (SWA) and standard MoE with no shared experts—marking a stark contrast with DeepSeek's approach. He notes the interesting parts are the data and experiments, skipping most architecture details.

Original post →

More from Research

Research channel →