Ollama preview powered by MLX boosts inference speed on Apple Silicon

awnihannun · x · 2026-08-16

Ollama released a preview powered by Apple's MLX framework, optimized for Apple Silicon. By leveraging unified memory architecture and GPU Neural Accelerators in M5 chips, the update significantly improves Time to First Token (TTFT) and generation speed. Tests with Qwen3.5-35B-A3B show prefill speeds reaching 1851 token/s and decode speeds of 134 token/s. The update also introduces NVFP4 support for higher quality inference and better memory efficiency.

Original post →

More from Infra

Infra channel →