Ollama preview powered by MLX boosts inference speed on Apple Silicon
awnihannun · x · 2026-08-16
Ollama released a preview powered by Apple's MLX framework, optimized for Apple Silicon. By leveraging unified memory architecture and GPU Neural Accelerators in M5 chips, the update significantly improves Time to First Token (TTFT) and generation speed. Tests with Qwen3.5-35B-A3B show prefill speeds reaching 1851 token/s and decode speeds of 134 token/s. The update also introduces NVFP4 support for higher quality inference and better memory efficiency.
More from Infra
- Expert Concerns Embedded Devices as a Weak Link in AI Security — johnowhitaker · 2026-08-17
- Help: Converting Qwen3.8 27B to ONNX for NPU Usage — xXDennisXx3000 · 2026-08-17
- Gonka offers 1M context V4-Flash via decentralized inference — autoimago · 2026-08-17
- Datacenter emissions are covering our glaciers — djcows · 2026-08-17
- Summary of Megakernel literature — marksaroufim · 2026-08-17
- Qwen3.8-27B hits 223 tok/s on RTX 6000 Pro with NVFP4 — BanghuaZ · 2026-08-17