MoVA explained: how IFM's 36B-A4B matches a dense 32B with 4B active params

rohanpaul_ai · x · 2026-09-11

K2 Horizon's efficiency comes from Mixture-of-Value Attention (MoVA), which extends MoE routing from FFN layers into attention itself.

Related event: IFM Open-Sources K2 Horizon: Six Models from 0.9B to 375B with Parallel Decoding and Benchmark Cheating Audit(7 posts)→

Original post →

More from Models

Models channel →