Leaked Architecture of ~400B MoE Model with Aggressive GQA Sparks Interest
teortaxesTex · x · 2026-08-10
A developer analyzed the architecture of a suspected new model, Muse Glimmer 30B. It is presumed to be a Mixture of Experts (MoE) model with around 400B total parameters and 15B active parameters. The architecture notably features a 3:1 local/global attention ratio paired with absurdly aggressive Grouped-Query Attention (GQA). Observers are amazed that such an extreme design works remarkably well.
More from Models
- Meta's New Open Multimodal Model Muse Glimmer Lands on Ollama — ollama · 2026-08-10
- Half the Models in the Third Tier Are Completely Lost — teortaxesTex · 2026-08-10
- Tencent Hunyuan Hy3 Hits #1 on OpenRouter, Offered Free via WorkBuddy — heyshrutimishra · 2026-08-10
- MOSS-TTS-Nano: Open-Source 0.1B Multilingual TTS Model Runs Realtime on CPU — tom_doerr · 2026-08-10
- Speculative Decoding Boosts RTX 5090 to 233 tok/s, Outperforming Mac — rohanpaul_ai · 2026-08-10
- Hardcore Reverse Engineering: Developer Rebuilds Kimi K3 Training Pipeline from Scratch — sharpeye_wnl · 2026-08-10