MLX-VLM tops Apple Silicon inference speed with 36.4 tok/s, beating Ollama and llama.cpp
andrejusb · x · 2026-08-16
MLX-VLM is touted as the most mature and fastest inference engine for Apple Silicon. In tests on M3 Max 96GB, it hit 36.4 tok/s, surpassing Ollama (29.0), MTPLX (25.3), and others. It powers Nativ AI and has 200+ contributors.
More from Infra
- AI's Next Bottleneck Is Power — ingliguori · 2026-08-16
- 1-bit Quantized Qwen 27B Runs on 12GB VRAM at 92 tokens/s — zyxciss · 2026-08-16
- Google's HEIR Compiler Ships Demos for Encrypted AI Inference — heypearlai · 2026-08-16
- GLM 5.3 Update Pace Sparks Compute Comparison with DeepSeek — teortaxesTex · 2026-08-16
- Alibaba launches Zhenwu M890 SuperNode, supports 122k-card clusters — pstAsiatech · 2026-08-16
- R1 runs at 150 tps, 7x faster than DeepSeek's serving — teortaxesTex · 2026-08-16