MoVA explained: how IFM's 36B-A4B matches a dense 32B with 4B active params
rohanpaul_ai · x · 2026-09-11
K2 Horizon's efficiency comes from Mixture-of-Value Attention (MoVA), which extends MoE routing from FFN layers into attention itself.
- The 36B model activates only 4B params per token yet lands close to the dense 32B
- Because the dense and sparse models were trained under identical conditions, they offer a clean dense-vs-sparse comparison instead of comparing unrelated models
More from Models
- User: Opus 5 is "unusable" — it finds every way not to do what's asked even at max settings — RexDouglass · 2026-09-11
- Only Muse Spark 1.3 and Fable 5.1 sit on the coding Pareto frontier — jyangballin · 2026-09-11
- User slams Anthropic for blocking benign queries on ancient texts and recursive AI — NickPassig · 2026-09-11
- Codex cybersecurity work needs the Daybreak model to avoid safety-guardrail blocks — HankYeomans · 2026-09-11
- DeepSeek's new open-source model reportedly crushes GLM and Kimi at 4-10x lower prices — anselm · 2026-09-11
- GPT-6 Astra rebuilds Cessna 337 landing gear from a YouTube video; Fable 5.1 falls short — FinanceYF5 · 2026-09-11