Post says an attention-heavy model has 104B active parameters
zephyr_z9 · x · 2026-07-28
A short post quotes a note about a "very heavy attention backbone" and says the model has 104B active parameters, much larger than expected.
With no further context, the key signal is the scale claim itself: the architecture appears to use a much heavier active subset than the commenter anticipated.
More from Models
- Kimi K3’s MoE routing may be driving higher expert-parallel communication costs — stochasticchasm · 2026-07-28
- Kimi K3 finds 16 new vulnerabilities and beats GLM-5.2 on an exploit benchmark — zephyr_z9 · 2026-07-28
- Kimi K3 weight shard appears as `model-00001-of-000096.safetensors` — ricklamers · 2026-07-28
- Microsoft launches MAI-Cyber-1-Flash and MDASH, claiming top CyberGym results at half the cost — satyanadella · 2026-07-28
- Claude Opus 5’s migration guide quietly changes years of prompting advice — AlexKim · 2026-07-28
- MiMo V2.5 Pro beats DeepSeek V4 Pro in a World Cup prediction test — PreciousSeige · 2026-07-28