Apple Silicon inference runs 137% faster as MLX Challenge pushes edge limits

gajesh · x · 2026-08-02

The MLX Fast Challenge has demonstrated the massive untapped potential of Apple Silicon for local LLM inference. The current 1st place entry achieves a 137% speedup over the initial baseline, hitting 184.6 decode tok/s and 5078.6 prefill tok/s, compared to the baseline's 73.7 and 2616.4 respectively.

Original post →

More from Infra

Infra channel →