First community MLX 4-bit benchmarks for K2-Horizon-MoVA-36B hit 49.1 tok/s locally

DerTomsn · reddit · 2026-09-04

First community benchmark results for K2-Horizon-MoVA-36B-A4B on oMLX (MLX 4-bit) are live on llm-bench.io, reaching up to 49.1 tok/s for local inference. The author notes it runs a bit slower than existing A3B-class models, likely due to missing MTP (multi-token prediction), and links an A3B oQ8e comparison for context.

Original post →

More from Infra

Infra channel →