K2 Horizon Lineup Hits Artificial Analysis, but Awful KV Cache Design Misleads Its Charts

crusaderky · reddit · 2026-09-14

The full K2 Horizon lineup is on Artificial Analysis (3.7B/7B SOTA, 36B-A4B good for low-bandwidth hardware), but the author argues its terrible KV cache design makes parameter counts misleading: with 128k context the 3.7B uses 5 GiB for context vs Qwen3.6-35B-A3B's 0.7 GiB. Practical picks: 36B-A4B only on exactly 16GB VRAM + ≥32GB RAM; on 24GB, Qwen3.8-27B is faster, smarter, and supports 256k context.

Original post →

More from Infra

Infra channel →