Apple Silicon inference runs 137% faster as MLX Challenge pushes edge limits
gajesh · x · 2026-08-02
The MLX Fast Challenge has demonstrated the massive untapped potential of Apple Silicon for local LLM inference. The current 1st place entry achieves a 137% speedup over the initial baseline, hitting 184.6 decode tok/s and 5078.6 prefill tok/s, compared to the baseline's 73.7 and 2616.4 respectively.
More from Infra
- App Developers Should Ship Their Own On-Device Models — abacaj · 2026-08-03
- Nemotron 3 Nano Omni Hits 264 tok/s Native on DGX Spark — ivan_bezdomny · 2026-08-03
- Global AI compute to hit 200M H100-equivalents by 2028, fueling agentic loop toward ASI — 新智元 · 2026-08-03
- tinybox Dual-GPU Edition Hits 245 tok/s Running DeepSeek — AccBalanced · 2026-08-03
- Troubleshooting KV Cache Misses Caused by Multiple Agent Tool Calls — CentrifugalMalaise · 2026-08-03
- Wafer serves Kimi K3 on AMD MI355X with 3.8x throughput and 71% lower cost vs B200 — SumitGup · 2026-08-03