First community MLX 4-bit benchmarks for K2-Horizon-MoVA-36B hit 49.1 tok/s locally
DerTomsn · reddit · 2026-09-04
First community benchmark results for K2-Horizon-MoVA-36B-A4B on oMLX (MLX 4-bit) are live on llm-bench.io, reaching up to 49.1 tok/s for local inference. The author notes it runs a bit slower than existing A3B-class models, likely due to missing MTP (multi-token prediction), and links an A3B oQ8e comparison for context.
More from Infra
- LFM2-350M NVFP4A16 hits 1.7M tok/s decode on a single RTX 5090 — AlpinDale · 2026-09-04
- Lightning AI ships 200x faster Drive persistence, starts swapping H100s for H200e — LightningAI · 2026-09-04
- LLM inference metrics explained: what TTFT, TPS, TPOT actually measure — abhijithneil · 2026-09-04
- LLM decoding explained: prefill reads in parallel, decode writes token by token — abhijithneil · 2026-09-04
- GPU Inference Explained: Memory Bandwidth, Not Compute, Caps Tokens Per Second — abhijithneil · 2026-09-04
- LLM inference 101: bandwidth ÷ weight bytes gives your throughput ceiling before any code — abhijithneil · 2026-09-04