K2 Horizon Lineup Hits Artificial Analysis, but Awful KV Cache Design Misleads Its Charts
crusaderky · reddit · 2026-09-14
The full K2 Horizon lineup is on Artificial Analysis (3.7B/7B SOTA, 36B-A4B good for low-bandwidth hardware), but the author argues its terrible KV cache design makes parameter counts misleading: with 128k context the 3.7B uses 5 GiB for context vs Qwen3.6-35B-A3B's 0.7 GiB. Practical picks: 36B-A4B only on exactly 16GB VRAM + ≥32GB RAM; on 24GB, Qwen3.8-27B is faster, smarter, and supports 256k context.
More from Infra
- PufferLib 5.0 Hits 60M Steps/Sec on a Single RTX 5090: How 5 Degrees of Parallelism Work — jsuarez · 2026-09-14
- Running a 510GB DeepSeek model on a 128GB DGX Spark by pruning unused MoE experts — pbaylies · 2026-09-14
- Meta's ZGateway proxy cut ZippyDB per-host connections by 97-98% with just 6% compute overhead — Meta_Engineers · 2026-09-14
- PufferLib 5.0 is ~5x faster than 4.0, peak throughput over 60M SPS — jsuarez · 2026-09-14
- PufferLib 5.0 ditches Python: bitwise-deterministic async training in under 10k lines of CUDA C — jsuarez · 2026-09-14
- Baseten's Philip Kiely releases 'Inference Engineering' book free online — philipkiely · 2026-09-14