IFM's K2-Horizon-36B-A4B Matches 20x-Larger Models on AA Index Using New MoVA Architecture
victormustar · x · 2026-09-22
IFM has open-sourced K2-Horizon-36B-A4B, which scores 25 on the Artificial Analysis Intelligence Index while activating only 4B parameters per token — matching models with over 20x its total parameters.
The gains come from MoVA (Mixture-of-Value Attention), a new architecture that applies MoE-style sparsity to the value vectors in multi-head attention, opening a second sparsity scaling axis beyond FFN-level MoE:
- Simple and compatible with efficient attention methods (FlashAttention, GQA, sparse attention)
- No extra KV cache cost compared to standard GQA
The model is available on Hugging Face as IFM/K2-Horizon-MoVA-36B-A4B.
More from Models
- Reliquary-4B: A 4B math & code model trained via decentralized RL with community rollouts — const_reborn · 2026-09-22
- Users say they can't trick Jev into hallucinating — BLUECOW009 · 2026-09-22
- Measured trade-offs of three REAP-pruned Qwen3.8-Flash-Next MLX builds on Apple Silicon — MensaProdigy · 2026-09-22
- Dev claims further-optimized DeepSeek V4 NVFP4 uses 190GB of 192GB VRAM — HankYeomans · 2026-09-22
- OpenAI researcher Will Depue on why voice models still lack true realtime chat — willdepue · 2026-09-22
- Hands-on: Ling-3.0-flash-VL generates full web page code from a screenshot in 17 seconds — _jaydeepkarale · 2026-09-22