MLX-VLM tops Apple Silicon inference speed with 36.4 tok/s, beating Ollama and llama.cpp

andrejusb · x · 2026-08-16

MLX-VLM is touted as the most mature and fastest inference engine for Apple Silicon. In tests on M3 Max 96GB, it hit 36.4 tok/s, surpassing Ollama (29.0), MTPLX (25.3), and others. It powers Nativ AI and has 200+ contributors.

Original post →

More from Infra

Infra channel →