Deep Dive into a Model's Attention and MoE Details
nrehiew_ · x · 2026-07-16
This post breaks down a model's architectural design, focusing on two main areas:
Attention Design
- Uses sliding window, which is somewhat surprising given the team's strong background in linear attention.
- Applies basic 1D convolution on KV and residual.
- No RoPE is used; instead, attention weights are formed by projecting the hidden state and combining it with distance-dependent position bias.
MoE Design
- The model scale is about 975B parameters with 41B active, trained on over 45T multimodal tokens.
- Compared to Dsv3, active parameters increased by roughly 10%, and token count is 3 times larger.
- Features 6 routed experts out of 256 experts, plus 2 shared experts; the author notes that having 2 shared experts is quite rare, as the norm is typically 0 or 1.
Overall, a highly detailed architectural observation post.
Related event: Inkling Performance and Related Architecture Speculation(7 posts)→
More from Models
- Moonshot pauses Kimi K3 signups five days after launch as GPU demand surges — eyishazyer · 2026-07-21
- AI Diplomacy demo makes agents negotiate, ally, and betray each other — jamdac · 2026-07-21
- Newer models need a different prompting style, and old tricks can make outputs worse — emollick · 2026-07-21
- GLM-5.5 is said to arrive in 4 weeks with open weights — tanay_mehta · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Ben’s Bites roundup highlights Kimi K3, Fable 5, Cursor costs and self-driving companies — Ben's Bites · 2026-07-21