Suspected Frontier Model Architecture: No RoPE and Hybrid Attention

nrehiew_ · x · 2026-07-16

The post analyzes the architectural details of a suspected frontier model. Regarding the attention mechanism: it adopts a sliding window and uses 1D convolution on KV and residual connections; RoPE is completely abandoned in favor of distance-based relative position bias. On the MoE side: total parameters are 975B with 41B active, trained on 45T multimodal tokens; it uses an architecture similar to DeepSeek v3 but features 6 routed experts and 2 shared experts (usually 1 or 0), thereby increasing the sparsity rate.

Related event: Inkling Performance and Related Architecture Speculation(7 posts)→

Original post →

More from Research

Research channel →