Leaked Architecture of ~400B MoE Model with Aggressive GQA Sparks Interest

teortaxesTex · x · 2026-08-10

A developer analyzed the architecture of a suspected new model, Muse Glimmer 30B. It is presumed to be a Mixture of Experts (MoE) model with around 400B total parameters and 15B active parameters. The architecture notably features a 3:1 local/global attention ratio paired with absurdly aggressive Grouped-Query Attention (GQA). Observers are amazed that such an extreme design works remarkably well.

Original post →

More from Models

Models channel →